datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ne-tts-coqui-multilingual
NE-TTS Coqui Multilingual
Multilingual TTS dataset for 15 North East Indian languages, formatted for Coqui-AI VITS multilingual training. Contains 61,943 clips / 83.5 hours at 22050Hz (SNR >= 20dB only).
Languages
ISO
Language
Clips
Hours
grt
Garo
24,772
29.6
ccp
Chakma
10,689
14.3
nag
Nagamese
9,688
14.5
lus
Mizo
8,554
14.5
nnp
Wancho
5,081
6.3
trp
Kokborok
1,237
1.7
clk
Idu Mishmi
602
0.7
mjw
Karbi
373
0.4
nre
Rengma
258
0.4
nri… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-coqui-multilingual.coqar-clarifications-audio
CoQAR Clarifications with synthetic context audio
These audio recordings are AI-generated speech, not recordings of human speakers.
OpenAI tts-1 narrated each exact story using voice alloy, speed 1,
and MP3 output. Long stories are synthesized in ordered parts and joined; see the
audio generation manifest for part boundaries and measured audio properties.
No questions, answers, rationales, or stored model prompts were narrated.
The original appended and inserted configurations… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/coqar-clarifications-audio.missing_FR_drugs_and_units_with_coqui
