CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abnajlae /darija-asr-corpus Darija ASR Corpus (dataset-core) Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). This repo contains four source subsets: DODa, DVoice, Wiki, and YouTube. Each subset carries its own upstream license/terms -- see below -- because they are drawn from four different original projects. Subsets Config Rows Audio bundled? Upstream license Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.audioautomatic-speech-recognition10K<n<100K0 likes473 downloads16d agoHugging Face02XAILab-CyberSpark /Cabin-Human-ABNORMAL-Behavior-Dataset 全球最大的智能座舱多模态开源高质量数据集来啦! 一. 数据集摘要 (Dataset Summary) 「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。 该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。 数据格式 数据集以JSON格式提供,包含以下字段: image_id: 图像ID image_path: 图像路径 category: 行为类别 tags: 行为标签 behaviors: 包含左右乘客行为描述的对象 left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-ABNORMAL-Behavior-Dataset.image1K<n<10K0 likes280 downloads1y agoHugging Face03Abner0803 /amzn_synthetic_conversation_title_id-aligntext1K<n<10K0 likes120 downloads10mo agoHugging Face04Abner0803 /Amazon-Beauty-S1This dataset is derived from Amazon Reviews'23 [1] Beauty category. The split is standard leave-one-out: the last item is the test target, the second-to-last is the validation target, and everything before that is training. The training portion is expanded by sliding window — every prefix becomes one example — so a user with a sequence of length $L$ contributes $L-3$ training rows with histories of length $1 \dots L-3$, one validation row and one test row. Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Abner0803/Amazon-Beauty-S1.texttext-generation10K<n<100K0 likes75 downloads19d agoHugging Face05Abner0803 /msmarco300k-rawtext100K<n<1M0 likes72 downloads16d agoHugging Face06abnajlae /darija-asr-benchmark-6speaker Darija ASR 6-Speaker Benchmark A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3 female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus), used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi) Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming). Consent and anonymization Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.audioautomatic-speech-recognitionn<1K0 likes67 downloads16d agoHugging Face07Abner0803 /nq320k-rawtext100K<n<1M0 likes51 downloads27d agoHugging Face08Abnik /spotdiff-v1-dev SpotDiff v1 Development Dataset SpotDiff is a visual spot-the-difference benchmark. Each item is one composite image containing two nearly identical panels. A model receives the complete image and identifies every visible difference using structured JSON. This is the public development release. It contains 10 reviewed images and 61 gold differences. The annotations are intentionally public so that anyone can reproduce the evaluator locally and inspect the benchmark design.… See the full description on the dataset page: https://huggingface.co/datasets/Abnik/spotdiff-v1-dev.textimage-to-textn<1K1 likes20 downloads1mo agoHugging Face09Abner0803 /amzn_conv_titleid_v2-14ktext10K<n<100K0 likes14 downloads9mo agoHugging Face10Abner0803 /amzn_title_14k_chattext10K<n<100K0 likes14 downloads9mo agoHugging Face11Abner0803 /amzn_semantic_conv_14ktext10K<n<100K0 likes14 downloads9mo agoHugging Face12Abner0803 /msmarco_stage1_title_100ktext100K<n<1M0 likes13 downloads8mo agoHugging Face13Abner0803 /msmarco-ICL-100k In-Context Learning Dataset for MSMARCO-100k train.jsonl Contains ~150k indexing (doc, docid) & retrieval (query, docid) pairs test.jsonl Contains ~20k unseen retrieval pairs icl_test.jsonl Contains ~10k unseen indexing & retrieval pairs textquestion-answering100K<n<1M0 likes13 downloads7mo agoHugging Face14Abner0803 /msmarco-icl-3shot-no_copytext100K<n<1M0 likes12 downloads2mo agoHugging Face15Abner0803 /amzn_synthetic_conversation_semantic-id-align-docid_concattext10K<n<100K0 likes11 downloads10mo agoHugging Face16Abner0803 /amzn_conv_titleid_v2text1K<n<10K0 likes11 downloads9mo agoHugging Face17Abner0803 /amzn_title_5k_chattext10K<n<100K0 likes11 downloads9mo agoHugging Face18Abner0803 /amzn_stage1_title_100ktext100K<n<1M0 likes11 downloads9mo agoHugging Face19abneraigc /wiki_medical_terms_llamatext1K<n<10K0 likes10 downloads3y agoHugging Face20Abner0803 /amzn_synthetic_conversation_semantic-idtext1K<n<10K0 likes10 downloads10mo agoHugging Face21Abner0803 /amzn_title_id-14ktexttext-generation10K<n<100K0 likes10 downloads10mo agoHugging Face22happy8825 /abnormal_ecva_sfttext1K<n<10K0 likes9 downloads9mo agoHugging Face23Abner0803 /NQ-ICL-100ktextquestion-answering100K<n<1M0 likes9 downloads7mo agoHugging Face24Abner0803 /msmarco_text-with_pseudo_query-100k-grtext100K<n<1M0 likes9 downloads5mo agoHugging Face25Abner0803 /amzn_semantic_id-5ktext10K<n<100K0 likes8 downloads9mo agoHugging Face26Abner0803 /amzn_synthetic_conversation_semantic-id-aligntext1K<n<10K0 likes7 downloads10mo agoHugging Face27Abner0803 /msmarco-icl-100shot-id_onlytext100K<n<1M0 likes7 downloads4mo agoHugging Face28Abner0803 /amzn_semantic_idtext10K<n<100K0 likes6 downloads10mo agoHugging Face29dtldr /BGL_Abnormal_0205textn<1K0 likes6 downloads8mo agoHugging Face30Abner0803 /amzn_title_conv_5ktext1K<n<10K0 likes5 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.