datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ne-asr-dataset-lus-aug
NE ASR Augmented Dataset -- Mizo (lus)
Augmented automatic speech recognition dataset for Mizo (lus),
a Tibeto-Burman language spoken in Mizoram, India.
Source
Augmented from sulabhkatiyar/ne-asr-lus
(original transcribed speech data from the ARTPARK-IISc Vaani project).
Language Information
Property
Value
Language
Mizo
ISO 639-3
lus
Family
Tibeto-Burman
Region
Mizoram, India
Tonal
Yes
Tier
D (20.75h original data)… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-lus-aug.ne-asr-dataset-lus
Mizo (lus) — ASR dataset
A small Mizo (lus) speech-to-text dataset for automatic speech recognition
(ASR) of a low-resource North-East India language. Each example pairs a short audio
clip with its Romanized (Latin-script) transcript.
Source
Derived from the ARTPARK-IISc Vaani project (https://vaani.iisc.ac.in/)
Splits
Split
Samples
train
9,850
validation
1,190
test
1,218
Data fields
Each example has:
audio — the… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-asr-dataset-lus.ne-tts-lus
NE-TTS Mizo (lus)
Cleaned TTS dataset for Mizo (lus), a North East Indian language. Derived from the Vaani dataset with SNR filtering, LUFS normalization, and text cleaning.
Stats
Metric
Value
Total clips
11,766
Total hours
19.9h
High SNR (>=20dB)
8,554 clips
Medium SNR (15-20dB)
3,212 clips
Sample rate
16kHz
Audio format
WAV, 16-bit PCM
Schema
Column
Type
Description
audio
Audio
16kHz WAV audio
text
string… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-lus.ne-tts-f5-lus
NE-TTS F5 Mizo (lus)
F5-TTS training dataset for Mizo (lus). Contains 8,554 clips at 24kHz (SNR >= 20dB) from the cleaned NE-TTS dataset, formatted for F5-TTS training.
Stats
Metric
Value
Clips
8,554
Hours
14.5h
Sample rate
24kHz
SNR filter
>= 20dB
Source
ne-tts-lus
Schema
Column
Type
Description
audio
Audio
24kHz WAV audio
text
string
Cleaned transcript
Note
Audio is upsampled from 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-f5-lus.Luster-Flixvozfolha
