CoolFace
Datasetpublic

projectkaira/Pretraining-V1

Indic TTS Unified v1 A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio. All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/Pretraining-V1.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes12kdownloads

projectkaira/Pretraining-V1 · main · files are served by the source, never re-hosted here

projectkaira/Pretraining-V1 · CoolFace