sulabhkatiyar/ne-asr-dataset-lus
Mizo (lus) — ASR dataset A small Mizo (lus) speech-to-text dataset for automatic speech recognition (ASR) of a low-resource North-East India language. Each example pairs a short audio clip with its Romanized (Latin-script) transcript. Source Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/) Splits Split Samples train 9,850 validation 1,190 test 1,218 Data fields Each example has: audio — the… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-lus.
docs: rename + Vaani attribution
docs: dataset card
Add dataset card for Mizo (lus) ASR dataset
Add test shards 1-3/3 for Mizo (lus)
Add validation shards 1-3/3 for Mizo (lus)
Add train shards 16-20/20 for Mizo (lus)
Add train shards 11-15/20 for Mizo (lus)
Add train shards 6-10/20 for Mizo (lus)
Add train shards 1-5/20 for Mizo (lus)
initial commit
