CoolFace
Datasetpublic

sulabhkatiyar/ne-asr-dataset-lus

Mizo (lus) — ASR dataset A small Mizo (lus) speech-to-text dataset for automatic speech recognition (ASR) of a low-resource North-East India language. Each example pairs a short audio clip with its Romanized (Latin-script) transcript. Source Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/) Splits Split Samples train 9,850 validation 1,190 test 1,218 Data fields Each example has: audio — the… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-lus.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
1likes386downloads
10 commits on main
76819ef1mo ago

docs: rename + Vaani attribution

sulabhkatiyar
ae70b9c1mo ago

docs: dataset card

sulabhkatiyar
60704594mo ago

Add dataset card for Mizo (lus) ASR dataset

sulabhkatiyar
fc9ce634mo ago

Add test shards 1-3/3 for Mizo (lus)

sulabhkatiyar
f9ec81f4mo ago

Add validation shards 1-3/3 for Mizo (lus)

sulabhkatiyar
eef26ce4mo ago

Add train shards 16-20/20 for Mizo (lus)

sulabhkatiyar
edc822d4mo ago

Add train shards 11-15/20 for Mizo (lus)

sulabhkatiyar
7a8aa474mo ago

Add train shards 6-10/20 for Mizo (lus)

sulabhkatiyar
0b5d78f4mo ago

Add train shards 1-5/20 for Mizo (lus)

sulabhkatiyar
e2dc8444mo ago

initial commit

sulabhkatiyar