datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
darija-asr-corpus
Darija ASR Corpus (dataset-core)
Arabizi (Latin-script) transcriptions of Moroccan Darija speech, produced for a
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
This repo contains four source subsets: DODa, DVoice, Wiki, and
YouTube. Each subset carries its own upstream license/terms -- see below --
because they are drawn from four different original projects.
Subsets
Config
Rows
Audio bundled?
Upstream license
Upstream source… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-corpus.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-ABNORMAL-Behavior-Dataset.amzn_synthetic_conversation_title_id-alignAmazon-Beauty-S1This dataset is derived from Amazon Reviews'23 [1] Beauty category. The split is standard leave-one-out: the last item is the test target, the second-to-last is the validation target, and everything before that is training.
The training portion is expanded by sliding window — every prefix becomes one example — so a user with a sequence of length $L$ contributes $L-3$ training rows with histories of length $1 \dots L-3$, one validation row and one test row.
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Abner0803/Amazon-Beauty-S1.msmarco300k-rawdarija-asr-benchmark-6speaker
Darija ASR 6-Speaker Benchmark
A fixed, paired 20-utterance benchmark read identically by 6 held-out speakers (3
female: F1, F2, F3; 3 male: M1, M2, M3 -- none present in any training corpus),
used to evaluate cross-speaker generalization for a Moroccan Darija (Arabizi)
Whisper fine-tuning pipeline (paper not yet published -- citation forthcoming).
Consent and anonymization
Written informed consent was obtained from all six speakers for the recording and… See the full description on the dataset page: https://huggingface.co/datasets/abnajlae/darija-asr-benchmark-6speaker.nq320k-rawspotdiff-v1-dev
SpotDiff v1 Development Dataset
SpotDiff is a visual spot-the-difference benchmark. Each item is one composite
image containing two nearly identical panels. A model receives the complete
image and identifies every visible difference using structured JSON.
This is the public development release. It contains 10 reviewed images and 61
gold differences. The annotations are intentionally public so that anyone can
reproduce the evaluator locally and inspect the benchmark design.… See the full description on the dataset page: https://huggingface.co/datasets/Abnik/spotdiff-v1-dev.amzn_conv_titleid_v2-14kamzn_title_14k_chatamzn_semantic_conv_14kmsmarco_stage1_title_100kmsmarco-ICL-100k
In-Context Learning Dataset for MSMARCO-100k
train.jsonl
Contains ~150k indexing (doc, docid) & retrieval (query, docid) pairs
test.jsonl
Contains ~20k unseen retrieval pairs
icl_test.jsonl
Contains ~10k unseen indexing & retrieval pairs
msmarco-icl-3shot-no_copyamzn_synthetic_conversation_semantic-id-align-docid_concatamzn_conv_titleid_v2amzn_title_5k_chatamzn_stage1_title_100kwiki_medical_terms_llamaamzn_synthetic_conversation_semantic-idamzn_title_id-14kabnormal_ecva_sftNQ-ICL-100kmsmarco_text-with_pseudo_query-100k-gramzn_semantic_id-5kamzn_synthetic_conversation_semantic-id-alignmsmarco-icl-100shot-id_onlyamzn_semantic_idBGL_Abnormal_0205amzn_title_conv_5k
