datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WolneLektury-TTS-Polish
WolneLektury-TTS-Polish
A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors.
Dataset Statistics
Metric
Value
Total samples
383,710
Total duration
997 hours
Unique narrators
1207
Male samples
294,756 (767h)
Female samples
88,945 (230h)
Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.TTS-German
TTS-German
High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.
Processing Pipeline
Standardize → 24kHz mono WAV, loudness normalize
Transcribe → WhisperX word-level timestamps
Segment → ≤12s at word boundaries
Denoise → DeepFilterNet
Quality filter → DNSMOS ≥ 2.5
G2P → IPA phonemes (custom dictionary)
Statistics
Metric
Value
Samples
670,509
Hours
1250h
Sample rate
24kHz mono
Max duration
12s
Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.TTS-English-HiFiTTS-English-LibriTTSTTS-Romanian
TTS-Romanian
A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from CartiaAudio.eu — Romanian audiobooks.
Dataset Statistics
Metric
Value
Total samples
267,410
Total duration
720 hours
Unique speakers
456
Average duration
9.7 seconds
Average DNSMOS
3.84
Features
Field
Type
Description
__key__
string
Unique sample identifier
mp3
Audio
Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.TTS-Italian
TTS-Italian
A high-quality Italian speech dataset for text-to-speech and automatic speech recognition.
Data Sources
Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org.
Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.)
License: Public Domain
Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Italian.TTS-Hungarian
TTS-Hungarian
A large-scale, high-quality Hungarian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from MEK (Magyar Elektronikus Könyvtár) — Hungarian audiobooks.
Dataset Statistics
Metric
Value
Total samples
253,116
Total duration
702 hours
Unique speakers
100
Average duration
10.0 seconds
Average DNSMOS
3.68
Features
Field
Type
Description
__key__
string
Unique sample… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Hungarian.TTS-Swedish
TTS-Swedish
A high-quality Swedish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from LibriVox — Swedish audiobooks.
Dataset Statistics
Metric
Value
Total samples
14,535
Total duration
40 hours
Unique speakers
9
Average duration
10.0 seconds
Average DNSMOS
3.69
Gender Distribution
Gender
Samples
Hours
Male
11,219
31.2
Female
3,316
9.3
Features
Field… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Swedish.TTS-Finnish
TTS-Finnish
A high-quality Finnish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from LibriVox — Finnish audiobooks.
Dataset Statistics
Metric
Value
Total samples
11,483
Total duration
32 hours
Unique speakers
8
Average duration
9.9 seconds
Average DNSMOS
3.84
Gender Distribution
Gender
Samples
Hours
Male
2,562
7.1
Female
8,921
24.4
Features
Field… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Finnish.TTS-Greek
TTS-Greek
A large-scale, high-quality Greek speech dataset for text-to-speech and automatic speech recognition.
Data Sources
This dataset combines two sources:
Source
Samples
Hours
License
Content
LibriVox
34,727
96.8
Public Domain
Modern Greek classic literature, philosophy, fiction
FLEURS-R (Google)
4,124
12.6
CC-BY 4.0
Wikipedia-sourced sentences, AI-restored audio
Dataset Statistics
Metric
Value
Total samples
38,851
Total… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Greek.TTS-Danish
TTS-Danish
A large-scale, high-quality Danish speech dataset for text-to-speech and automatic speech recognition.
Data Sources
This dataset combines three sources:
Source
Samples
Hours
License
Content
lydbog.com
35,719
92.0
CC-BY-SA 4.0
Danish classic literature, read by Kristoffer Hunsdahl
CoRal-TTS (Alexandra Institute)
19,996
30.5
CC0
Professional TTS recordings, 2 speakers
LibriVox
0
0.0
Public Domain
Danish audiobooks
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Danish.TTS-DutchParler-TTS-Datadriven-100h-44.1kHz_stage1TTS-Polish-GosiaTTS-Polish-NemoTTS-Polish-McTTS-Polish-Darkman
