datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cwicr-construction-rates
CWICR — Construction Works, Items, Costs & Resources
A multilingual, machine-readable database of national construction rate books for 30 countries / language locales. Each rate is fully decomposed into its work composition and resource breakdown (labour, machinery, materials), with unit prices, hierarchical classification, and physical parameters preserved in the source language.
This dataset is the tabular source-of-truth behind the cwicr-vector-db-bgem3-v3 Qdrant snapshots. Use… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-construction-rates.WolneLektury-TTS-Polish
WolneLektury-TTS-Polish
A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors.
Dataset Statistics
Metric
Value
Total samples
383,710
Total duration
997 hours
Unique narrators
1207
Male samples
294,756 (767h)
Female samples
88,945 (230h)
Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.TTS-German
TTS-German
High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.
Processing Pipeline
Standardize → 24kHz mono WAV, loudness normalize
Transcribe → WhisperX word-level timestamps
Segment → ≤12s at word boundaries
Denoise → DeepFilterNet
Quality filter → DNSMOS ≥ 2.5
G2P → IPA phonemes (custom dictionary)
Statistics
Metric
Value
Samples
670,509
Hours
1250h
Sample rate
24kHz mono
Max duration
12s
Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.TTS-English-HiFiTTS-English-LibriTTSTTS-Romanian
TTS-Romanian
A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from CartiaAudio.eu — Romanian audiobooks.
Dataset Statistics
Metric
Value
Total samples
267,410
Total duration
720 hours
Unique speakers
456
Average duration
9.7 seconds
Average DNSMOS
3.84
Features
Field
Type
Description
__key__
string
Unique sample identifier
mp3
Audio
Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.TTS-Italian
TTS-Italian
A high-quality Italian speech dataset for text-to-speech and automatic speech recognition.
Data Sources
Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org.
Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.)
License: Public Domain
Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Italian.TTS-Hungarian
TTS-Hungarian
A large-scale, high-quality Hungarian speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from MEK (Magyar Elektronikus Könyvtár) — Hungarian audiobooks.
Dataset Statistics
Metric
Value
Total samples
253,116
Total duration
702 hours
Unique speakers
100
Average duration
10.0 seconds
Average DNSMOS
3.68
Features
Field
Type
Description
__key__
string
Unique sample… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Hungarian.TTS-Swedish
TTS-Swedish
A high-quality Swedish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from LibriVox — Swedish audiobooks.
Dataset Statistics
Metric
Value
Total samples
14,535
Total duration
40 hours
Unique speakers
9
Average duration
10.0 seconds
Average DNSMOS
3.69
Gender Distribution
Gender
Samples
Hours
Male
11,219
31.2
Female
3,316
9.3
Features
Field… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Swedish.TTS-Finnish
TTS-Finnish
A high-quality Finnish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from LibriVox — Finnish audiobooks.
Dataset Statistics
Metric
Value
Total samples
11,483
Total duration
32 hours
Unique speakers
8
Average duration
9.9 seconds
Average DNSMOS
3.84
Gender Distribution
Gender
Samples
Hours
Male
2,562
7.1
Female
8,921
24.4
Features
Field… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Finnish.TTS-Greek
TTS-Greek
A large-scale, high-quality Greek speech dataset for text-to-speech and automatic speech recognition.
Data Sources
This dataset combines two sources:
Source
Samples
Hours
License
Content
LibriVox
34,727
96.8
Public Domain
Modern Greek classic literature, philosophy, fiction
FLEURS-R (Google)
4,124
12.6
CC-BY 4.0
Wikipedia-sourced sentences, AI-restored audio
Dataset Statistics
Metric
Value
Total samples
38,851
Total… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Greek.TTS-Danish
TTS-Danish
A large-scale, high-quality Danish speech dataset for text-to-speech and automatic speech recognition.
Data Sources
This dataset combines three sources:
Source
Samples
Hours
License
Content
lydbog.com
35,719
92.0
CC-BY-SA 4.0
Danish classic literature, read by Kristoffer Hunsdahl
CoRal-TTS (Alexandra Institute)
19,996
30.5
CC0
Professional TTS recordings, 2 speakers
LibriVox
0
0.0
Public Domain
Danish audiobooks
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Danish.movie-genre-prediction
Dataset Card for Movie Genre Prediction
Link to Movie Genre Prediction Competition
By accessing this dataset, you accept the rules of the Movie Genre Prediction competition.
Organizer
Organizer of this competition is Data-Driven Science.
Join our FREE 3-Day Object Detection Challenge!
Email Usage
By accessing this dataset, you consent that your email will be used for communication purposes from Data-Driven Science.We do not share nor sell our mailing list.… See the full description on the dataset page: https://huggingface.co/datasets/datadrivenscience/movie-genre-prediction.TTS-DutchParler-TTS-Datadriven-100h-44.1kHz_stage1TTS-Polish-GosiaTTS-Polish-NemoTTS-Polish-McTTS-Polish-Darkman
