datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPRING_INX_Bengali_R2SPRING_INX_Punjabi_R2TTS-R24000-ThSPRING_INX_Tamil_R2SPRING_INX_Gujarati_R2SPRING_INX_Marathi_R2SPRING_INX_Malayalam_R2med_male_female_r2nepali-audio-reserve-r2
Nepali synthetic educational dialogue
~137.0 h of Nepali speech at 24 kHz. Two speakers per clip, 2-5 minutes, teacher/student turns.
A backup, not a release: the transcripts are machine-generated, and none of this
audio passed the quality gate that produced our training corpus.
Synthetic. Generated by a TTS model reading NCERT-style educational dialogue, code-mixed Nepali/English. No human speaker is recorded here.
Columns
column
contents
id
first 16… See the full description on the dataset page: https://huggingface.co/datasets/milanakdj/nepali-audio-reserve-r2.SPRING_INX_Assamese_R2r222
