datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bangali_local_dialect_ASR_HF_Dataset
BanglaMix — Code-Switching ASR in Bangladeshi Regional Dialects
BanglaMix is a speech dataset for Automatic Speech Recognition (ASR) on dialectal Bangladeshi Bengali mixed with English (code-switching). It covers 15 regional dialects and the natural Bengali–English code-switching common in informal Bangladeshi speech — a setting not covered by existing Bengali corpora, which address either dialects or code-switching, never both.
Clips
41,499 transcribed audio clips… See the full description on the dataset page: https://huggingface.co/datasets/niloycste68/Bangali_local_dialect_ASR_HF_Dataset.user_03aa5df890b64866be4aef51a01c0a8a_local_dataset
