datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kokborok
Speed-Tb Phase 1 Kokborok Narration
Dataset Description
The Kok Borok Speech Dataset, developed as part of the
Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over 200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script, making it
** one of the largest speech resources for the… See the full description on the dataset page: https://huggingface.co/datasets/speed-tb/kokborok.kokoro-dialogue-asr
Kokoro dialogue ASR set
500 single-speaker English clips of ~25-30 s, synthesized with
Kokoro-82M (voice af_heart) reading
procedurally generated spoken-monologue passages. Intended as augmentation for ASR
fine-tuning, not as a standalone training set.
split
clips
hours
train
406
3.14
test
94
0.72
Fields
audio — 16 kHz mono
transcription — orthographic transcript, exactly the text that was synthesized
topic — which passage template produced it… See the full description on the dataset page: https://huggingface.co/datasets/sachin6624/kokoro-dialogue-asr.
