datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
nepali-english-codemixed-asrnep_eng_code-mixed_asr_datasetcode-mixed-bangla-english-asrnep_eng_code-mixed_asr_datasetnepali-english-codemixed-asruk-en-code-mixed-asr-2h
uk-en-code-mixed-asr-2h
A 2-hour dataset of Ukrainian-English code-mixed speech for automatic speech recognition, recorded by a single male speaker across 16 speaking-style personas covering software engineering domains (.NET, React, DevOps, project management).
Samples: 448
Total duration: 2h 6m 35s
Language: 446 mixed (code-mixed) + 2 uk (monolingual)
Speaker: 1 male voice, 16 personas (varied pacing, formality, topic; 15 map to Valera, 1 to ValeraAlt)
Audio format: OGG Opus… See the full description on the dataset page: https://huggingface.co/datasets/vnikitin/uk-en-code-mixed-asr-2h.nep_eng_code-mixed_asr_dataset
