datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bodo
Speed-Tb Phase 1 Bodo Narration
Dataset Description
The Bodo Speech Dataset, developed as part of the
Speech Datasets and Models for Tibeto-Burman Languages (Project SpeeD-TB),
funded under Mission Bhashini, is a transcribed speech corpus of the language.
The full dataset comprises over 200 hours of high-quality audio recordings paired with accurate transcriptions in both IPA and Roman script, making it
** one of the largest speech resources for the language**… See the full description on the dataset page: https://huggingface.co/datasets/speed-tb/bodo.ASR-Bodo_5hrs
