datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ne-asr-dataset-nre
Rengma (nre) — ASR dataset
A small Rengma (nre) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
336
validation
11
test
28
Note on the test size. The test split has only 28 utterances… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nre.ne-asr-dataset-nri
Chakhesang (nri) — ASR dataset
A small Chakhesang (nri) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
261
validation
26
test
6
Note on the test size. The test split has only 6… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nri.ne-asr-dataset-nre-aug
NE ASR Augmented Dataset -- Rengma (nre)
Augmented automatic speech recognition dataset for Rengma (nre),
a Tibeto-Burman language spoken in Nagaland, India.
Source
Augmented from sulabhkatiyar/ne-asr-nre
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Rengma
ISO 639-3
nre
Family
Tibeto-Burman
Region
Nagaland, India
Tonal
Yes
Tier
B (0.51h original data)… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nre-aug.ne-asr-dataset-nri-aug
NE ASR Augmented Dataset -- Chakhesang (nri)
Augmented automatic speech recognition dataset for Chakhesang (nri),
a Tibeto-Burman language spoken in Nagaland, India.
Source
Augmented from sulabhkatiyar/ne-asr-nri
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Chakhesang
ISO 639-3
nri
Family
Tibeto-Burman
Region
Nagaland, India
Tonal
Yes
Tier
A (0.38h original… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-nri-aug.
