datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YouTube-Cantonese-Emilia
YouTube Cantonese — Emilia
2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by
running alvanlii/cantonese-youtube
through the Emilia
speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering).
Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn
label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in
both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.Whisper-Hallucination
Whisper Hallucination and Repetition Probes
This is a benchmark. Every evaluation config is test — do not fine-tune on it.
lexicon_synth is the exception: synthetic training material with its own train/test
split, and not one of the eight benchmark arms.
To build training data, exclude the items in
benchmark/exclusions.json
(546 FMA tracks, 1,168 FSD50K ids, 2,620 LibriSpeech utterances, the Malay stems). The
benchmark draws FSD50K eval and FMA shards 0–1, so training can use… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Whisper-Hallucination.voice-of-care-health-dataset
Voice of Care AI for Global Health Benchmark Dataset
Overview
This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research.
Dataset Summary
Property
Details
Language
Hausa
Modality
Audio + Text
Task(s)
e.g. Speech Recognition, Emotion Detection, Dialect Identification
Version
1.0.0
🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.
