datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dutch-tts-labeled-complete
Dutch TTS Dataset - Complete Labeled
A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data.
Quick Preview
The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config.
Dataset Description
This dataset contains Dutch speech recordings with rich metadata including:
Emotion labels (neutral, happy, sad, angry)
Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.complete-voiceai-speech-dataset
Silencio Voice AI Sample Dataset
Speaker-attributed spontaneous speech. 363 labelled contributors across 111 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment.
Hours
16.05
Clips
1,305
Speakers
363
Origin varieties
111
Languages
21
Configs
44
Speaker metadata
origin region / variety, mother tongue, gender, device, OS… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset.complete-voiceai-speech-dataset-anonymized
🎙️ Silencio Network: Voice AI Sample Dataset
📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages.
📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests.
🌍 Why Silencio Data?
Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you:
✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/complete-voiceai-speech-dataset-anonymized.
