CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chosenek /czech-speech-combined Czech Speech Combined Dataset Quality-filtered Czech speech dataset for TTS/ASR training. 146,153 clips across ~1,700 speakers from 5 sources. Sources Source Clips Hours Speakers Origin audiobooks 24,535 ~33h 13 Czech audiobook narrations audiobooks_new 28,131 ~39h 8+ Czech audiobook narrations yodas_czech 28,208 ~37h ~2,300 YODAS YouTube speech (quality-filtered) voxpopuli_czech 12,679 ~33h 45 VoxPopuli parliament speech commonvoice_czech 52… See the full description on the dataset page: https://huggingface.co/datasets/chosenek/czech-speech-combined.text100K<n<1M2 likes230 downloads4mo agoHugging Face02shunyalabs /czech-speech-datasetaudio1K<n<10K0 likes47 downloads1y agoHugging Face03Thomcles /Czech-Speech-Monospeaker-Honza Important This dataset comes from voxpopuli. We selected the most frequent male speaker in the dataset and created a separate single-speaker dataset. Processing performed: Recording of a neutral speaker, in large quantities Denoising with https://huggingface.co/speechbrain/sepformer-whamr16k audiotext-to-speech1K<n<10K1 likes23 downloads10mo agoHugging Face04adityarra07 /czech_train_data Dataset Card for "czech_train_data" More Information needed audio10K<n<100K1 likes16 downloads3y agoHugging Face05adityarra07 /czech_test Dataset Card for "czech_test" More Information needed audio1K<n<10K0 likes9 downloads3y agoHugging Face06Katya-Iukhn /audio_czechaudio1K<n<10K0 likes8 downloads2y agoHugging Face07Thomcles /YodaLingua-Czechgated YodaLingua-Czech YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Czech portion of the multilingual YodaLingua collection. 🧾 Dataset Overview Property Value Total clips 33,948 audio–transcription pairs Total duration 88 hours Speakers 3,618 distinct speakers Audio format MP3 • mono • 24 kHz •… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Czech.audiotext-to-speech10K<n<100K1 likes7 downloads3mo agoHugging Face08night12 /czech_time_shards_recordingsaudion<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.