datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pretraining-V1
Indic TTS Unified v1
A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio.
All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source… See the full description on the dataset page: https://huggingface.co/datasets/projectkaira/Pretraining-V1.Pretraining-V1
Indic TTS Unified v1
A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This dataset consolidates 17 distinct source datasets into a single, schema-normalized resource covering Indian / South Asian languages, plus major European, African, MENA, and Central Asian languages, with over 13.7 million utterances and 26,000+ hours of audio.
All audio is resampled to 24 kHz mono. Every row follows an identical schema regardless of source, enabling… See the full description on the dataset page: https://huggingface.co/datasets/neh7777/Pretraining-V1.audio_pre-training-v1.3OpenWhistle-Pretraining
OpenWhistle Pretraining Dataset
OpenWhistleNeurIPS26/OpenWhistle-Pretraining is the public unlabeled audio
dataset used for OpenWhistle pretraining. It contains 96 kHz dolphin acoustic
segments with timing and recording metadata, but no whistle/noise labels.
The main default config is the complete pretraining dataset. A smaller
deterministic review-sample config is also provided so reviewers can inspect
representative examples quickly without downloading the full dataset.… See the full description on the dataset page: https://huggingface.co/datasets/OpenWhistleNeurIPS26/OpenWhistle-Pretraining.DS_pretraining_1_100_hourpretraining
