datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
czech-speech-combined
Czech Speech Combined Dataset
Quality-filtered Czech speech dataset for TTS/ASR training.
146,153 clips across ~1,700 speakers from 5 sources.
Sources
Source
Clips
Hours
Speakers
Origin
audiobooks
24,535
~33h
13
Czech audiobook narrations
audiobooks_new
28,131
~39h
8+
Czech audiobook narrations
yodas_czech
28,208
~37h
~2,300
YODAS YouTube speech (quality-filtered)
voxpopuli_czech
12,679
~33h
45
VoxPopuli parliament speech
commonvoice_czech
52… See the full description on the dataset page: https://huggingface.co/datasets/chosenek/czech-speech-combined.czech-speech-datasetCzech-Speech-Monospeaker-Honza
Important
This dataset comes from voxpopuli.
We selected the most frequent male speaker in the dataset and created a separate single-speaker dataset.
Processing performed:
Recording of a neutral speaker, in large quantities
Denoising with https://huggingface.co/speechbrain/sepformer-whamr16k
czech_train_data
Dataset Card for "czech_train_data"
More Information needed
czech_test
Dataset Card for "czech_test"
More Information needed
audio_czechYodaLingua-Czech
YodaLingua-Czech
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Czech portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
33,948 audio–transcription pairs
Total duration
88 hours
Speakers
3,618 distinct speakers
Audio format
MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Czech.czech_time_shards_recordings
