CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DataDrivenConstruction /cwicr-construction-rates CWICR — Construction Works, Items, Costs & Resources A multilingual, machine-readable database of national construction rate books for 30 countries / language locales. Each rate is fully decomposed into its work composition and resource breakdown (labour, machinery, materials), with unit prices, hierarchical classification, and physical parameters preserved in the source language. This dataset is the tabular source-of-truth behind the cwicr-vector-db-bgem3-v3 Qdrant snapshots. Use… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-construction-rates.tabular10M<n<100M3 likes933 downloads5mo agoHugging Face02datadriven-company /WolneLektury-TTS-Polish WolneLektury-TTS-Polish A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors. Dataset Statistics Metric Value Total samples 383,710 Total duration 997 hours Unique narrators 1207 Male samples 294,756 (767h) Female samples 88,945 (230h) Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.audiotext-to-speech100K<n<1M2 likes904 downloads8mo agoHugging Face03datadriven-company /TTS-German TTS-German High-quality German speech dataset for TTS and ASR, derived from CML-TTS German. Processing Pipeline Standardize → 24kHz mono WAV, loudness normalize Transcribe → WhisperX word-level timestamps Segment → ≤12s at word boundaries Denoise → DeepFilterNet Quality filter → DNSMOS ≥ 2.5 G2P → IPA phonemes (custom dictionary) Statistics Metric Value Samples 670,509 Hours 1250h Sample rate 24kHz mono Max duration 12s Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.audiotext-to-speech1M<n<10M4 likes525 downloads7mo agoHugging Face04datadriven-company /TTS-English-HiFiaudio100K<n<1M1 likes358 downloads7mo agoHugging Face05datadriven-company /TTS-English-LibriTTSaudio100K<n<1M2 likes338 downloads7mo agoHugging Face06datadriven-company /TTS-Romanian TTS-Romanian A large-scale, high-quality Romanian speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from CartiaAudio.eu — Romanian audiobooks. Dataset Statistics Metric Value Total samples 267,410 Total duration 720 hours Unique speakers 456 Average duration 9.7 seconds Average DNSMOS 3.84 Features Field Type Description __key__ string Unique sample identifier mp3 Audio Audio… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Romanian.audiotext-to-speech100K<n<1M3 likes279 downloads7mo agoHugging Face07datadriven-company /TTS-Italian TTS-Italian A high-quality Italian speech dataset for text-to-speech and automatic speech recognition. Data Sources Derived from LibriVox Italian — volunteer-read Italian public domain audiobooks hosted on archive.org. Books: 23 Italian-language audiobooks (Dante, Pirandello, Verga, De Amicis, Collodi, Pascoli, etc.) License: Public Domain Processing: Standardized to 24kHz mono, WhisperX transcription (large-v3) with word-level alignment, segmented at word boundaries… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Italian.audiotext-to-speech10K<n<100K0 likes250 downloads7mo agoHugging Face08datadriven-company /TTS-Hungarian TTS-Hungarian A large-scale, high-quality Hungarian speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from MEK (Magyar Elektronikus Könyvtár) — Hungarian audiobooks. Dataset Statistics Metric Value Total samples 253,116 Total duration 702 hours Unique speakers 100 Average duration 10.0 seconds Average DNSMOS 3.68 Features Field Type Description __key__ string Unique sample… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Hungarian.audiotext-to-speech100K<n<1M1 likes181 downloads8mo agoHugging Face09datadriven-company /TTS-Swedish TTS-Swedish A high-quality Swedish speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from LibriVox — Swedish audiobooks. Dataset Statistics Metric Value Total samples 14,535 Total duration 40 hours Unique speakers 9 Average duration 10.0 seconds Average DNSMOS 3.69 Gender Distribution Gender Samples Hours Male 11,219 31.2 Female 3,316 9.3 Features Field… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Swedish.audiotext-to-speech10K<n<100K0 likes151 downloads8mo agoHugging Face10datadriven-company /TTS-Finnish TTS-Finnish A high-quality Finnish speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from LibriVox — Finnish audiobooks. Dataset Statistics Metric Value Total samples 11,483 Total duration 32 hours Unique speakers 8 Average duration 9.9 seconds Average DNSMOS 3.84 Gender Distribution Gender Samples Hours Male 2,562 7.1 Female 8,921 24.4 Features Field… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Finnish.audiotext-to-speech10K<n<100K0 likes96 downloads8mo agoHugging Face11datadriven-company /TTS-Greek TTS-Greek A large-scale, high-quality Greek speech dataset for text-to-speech and automatic speech recognition. Data Sources This dataset combines two sources: Source Samples Hours License Content LibriVox 34,727 96.8 Public Domain Modern Greek classic literature, philosophy, fiction FLEURS-R (Google) 4,124 12.6 CC-BY 4.0 Wikipedia-sourced sentences, AI-restored audio Dataset Statistics Metric Value Total samples 38,851 Total… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Greek.audiotext-to-speech10K<n<100K0 likes84 downloads8mo agoHugging Face12datadriven-company /TTS-Danish TTS-Danish A large-scale, high-quality Danish speech dataset for text-to-speech and automatic speech recognition. Data Sources This dataset combines three sources: Source Samples Hours License Content lydbog.com 35,719 92.0 CC-BY-SA 4.0 Danish classic literature, read by Kristoffer Hunsdahl CoRal-TTS (Alexandra Institute) 19,996 30.5 CC0 Professional TTS recordings, 2 speakers LibriVox 0 0.0 Public Domain Danish audiobooks Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-Danish.audiotext-to-speech10K<n<100K0 likes47 downloads8mo agoHugging Face13datadrivenscience /movie-genre-predictiongated Dataset Card for Movie Genre Prediction Link to Movie Genre Prediction Competition By accessing this dataset, you accept the rules of the Movie Genre Prediction competition. Organizer Organizer of this competition is Data-Driven Science. Join our FREE 3-Day Object Detection Challenge! Email Usage By accessing this dataset, you consent that your email will be used for communication purposes from Data-Driven Science.We do not share nor sell our mailing list.… See the full description on the dataset page: https://huggingface.co/datasets/datadrivenscience/movie-genre-prediction.text10K<n<100K11 likes46 downloads3y agoHugging Face14datadriven-company /TTS-Dutchaudio100K<n<1M1 likes29 downloads7mo agoHugging Face15Irina25P /Parler-TTS-Datadriven-100h-44.1kHz_stage1audio10K<n<100K0 likes26 downloads4mo agoHugging Face16datadriven-company /TTS-Polish-Gosiaaudio1K<n<10K0 likes17 downloads7mo agoHugging Face17datadriven-company /TTS-Polish-Nemoaudio1K<n<10K1 likes13 downloads7mo agoHugging Face18datadriven-company /TTS-Polish-Mcaudio10K<n<100K0 likes11 downloads7mo agoHugging Face19datadriven-company /TTS-Polish-Darkmanaudio1K<n<10K0 likes4 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.