datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ne-asr-dataset-lus-aug
NE ASR Augmented Dataset -- Mizo (lus)
Augmented automatic speech recognition dataset for Mizo (lus),
a Tibeto-Burman language spoken in Mizoram, India.
Source
Augmented from sulabhkatiyar/ne-asr-lus
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Mizo
ISO 639-3
lus
Family
Tibeto-Burman
Region
Mizoram, India
Tonal
Yes
Tier
D (20.75h original data)… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-lus-aug.ne-asr-dataset-lus
Mizo (lus) — ASR dataset
A small Mizo (lus) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
9,850
validation
1,190
test
1,218
Data fields
Each example has:
audio — the… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-lus.
