CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01parler-tts /libritts_r_filtered Dataset Card for Filtered LibriTTS-R This is a filtered version of LibriTTS-R. It has been filtered based on two sources: LibriTTS-R paper [1], which lists samples for which speech restoration have failed LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected. LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.audiotext-to-speech100K<n<1M24 likes4.7k downloads2y agoHugging Face02TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.1k downloads2y agoHugging Face03kadirnar /voicehub-arena-seed-tts-eval VoiceHub Arena — full English Seed-TTS-Eval 35,904 synthesized WAV files: 33 model families × the same 1,088 target texts. The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB. All 198 shards and every WAV SHA256 were verified after backup. Interactive leaderboard and all audio samples · Source repository (access required). Contents audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.audiotext-to-speech10K<n<100K0 likes2.8k downloads8d agoHugging Face04multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.8k downloads3mo agoHugging Face05parler-tts /mls_eng Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.audioautomatic-speech-recognition10M<n<100M40 likes2.4k downloads2y agoHugging Face06RidheshBhati /tts_farm Multilingual TTS/ASR Aggregated Dataset Cleaned, deduplicated and loudness-normalized Arabic, Japanese, Korean, Turkish, and Vietnamese speech. The training columns are audio (16-bit PCM WAV, 22050 Hz) and text; the remaining columns contain quality and provenance metadata. audiotext-to-speech100K<n<1M2 likes2.2k downloads2mo agoHugging Face07TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes2.1k downloads6mo agoHugging Face08ghanaopenai /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes1.5k downloads2mo agoHugging Face09parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.2k downloads2y agoHugging Face10TTS-AGI /commonvoice22-sidon-dacvae CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents Source sarulab-speech/commonvoice22_sidon Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.audioautomatic-speech-recognition1M<n<10M1 likes1.1k downloads6mo agoHugging Face11ghanaopenai /ewe-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes892 downloads3mo agoHugging Face12datadriven-company /WolneLektury-TTS-Polish WolneLektury-TTS-Polish A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors. Dataset Statistics Metric Value Total samples 383,710 Total duration 997 hours Unique narrators 1207 Male samples 294,756 (767h) Female samples 88,945 (230h) Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.audiotext-to-speech100K<n<1M2 likes881 downloads8mo agoHugging Face13ShiniChien /TTSDistil-Phonologygated Overview This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training. The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization. This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil-Phonology.audiotext-to-speech100K<n<1M1 likes857 downloads2mo agoHugging Face14alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes804 downloads6mo agoHugging Face15pnnbao-ump /VieNeu-TTS-140hgated pnnbao-ump/VieNeu-TTS-140h Mô tả Dataset A high-quality Vietnamese Text-to-Speech (TTS) dataset containing 74,858 audio samples with phonemized transcripts. This benchmark dataset is designed for fine-tuning modern TTS models with maximum synthesis quality. The text corpus is completely phonemized using standard international phonetic alphabet (IPA) representations suitable for neural acoustic modeling. Quick Facts Language: Vietnamese 🇻🇳 Tasks:… See the full description on the dataset page: https://huggingface.co/datasets/pnnbao-ump/VieNeu-TTS-140h.audiotext-to-speech10K<n<100K35 likes780 downloads11d agoHugging Face16ghanaopenai /ghana-named-entities-tts-twi This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghana Named Entities TTS — Twi A Twi-language speech dataset built from descriptions of Ghana named entities (people, places, organisations, and concepts). Each audio clip is a synthesised reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.audiotext-to-speech1K<n<10K0 likes711 downloads3mo agoHugging Face17AITRADER /dutch-tts-labeled-complete Dutch TTS Dataset - Complete Labeled A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data. Quick Preview The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config. Dataset Description This dataset contains Dutch speech recordings with rich metadata including: Emotion labels (neutral, happy, sad, angry) Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.audiotext-to-speech100K<n<1M0 likes698 downloads9mo agoHugging Face18ghanaopenai /dagbani-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 53410 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes631 downloads3mo agoHugging Face19ShiniChien /TTSDistil2gated Overview This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training. The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization. This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil2.audiotext-to-speech100K<n<1M1 likes570 downloads13d agoHugging Face20datadriven-company /TTS-German TTS-German High-quality German speech dataset for TTS and ASR, derived from CML-TTS German. Processing Pipeline Standardize → 24kHz mono WAV, loudness normalize Transcribe → WhisperX word-level timestamps Segment → ≤12s at word boundaries Denoise → DeepFilterNet Quality filter → DNSMOS ≥ 2.5 G2P → IPA phonemes (custom dictionary) Statistics Metric Value Samples 670,509 Hours 1250h Sample rate 24kHz mono Max duration 12s Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.audiotext-to-speech1M<n<10M4 likes532 downloads7mo agoHugging Face211rsh /tts-rj-hi-karya Rajasthani Hindi Speech Dataset This dataset consists of audio recordings of participants reading out stories in Rajasthani Hindi, one sentence at a time. They had 98 participants from Soda, Rajasthan. Each participant read 30 stories. In total, we have 426872 recordings in this dataset. They had roughly 58 male participants and 40 female participants. Point to Note: While random sampling suggests that most users have to their best effort tried to accurately read out the sentences… See the full description on the dataset page: https://huggingface.co/datasets/1rsh/tts-rj-hi-karya.audiotext-to-speech100K<n<1M4 likes452 downloads3y agoHugging Face22ghananlpcommunity /twi-tts-asr Ghana Twi Speech Dataset (TTS + ASR) 79,655 synthetic Twi speech samples (88.8 hours) generated with OmniVoice. Designed for both text-to-speech (TTS) and automatic speech recognition (ASR) research on Twi. Intended use TTS / Voice cloning: Use the audio + text pairs directly. The dataset includes diverse voice profiles for building or fine-tuning Twi TTS systems. ASR: Use the text transcripts as labels. The audio is natively at 24 kHz — resample to 16 kHz for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-tts-asr.audiotext-to-speech10K<n<100K0 likes449 downloads2mo agoHugging Face23Anilosan15 /Turkish_TTS_Dataaudiotext-to-speech10K<n<100K21 likes443 downloads7mo agoHugging Face24Chalermdej /yodas2_sidon_th_tts Thai TTS Dataset — Filtered & Quality-Verified from YODAS2 sidon A filtered, quality-verified Thai text-to-speech dataset derived from sarulab-speech/yodas2_sidon, with transcriptions verified by multiple ASR models and Gemini, text fully normalized to Thai, and audio quality-screened with DNSMOS. Dataset Summary Samples 141,927 Audio hours 156.0 Speakers 4,199 Sample rate 24,000 Hz Format WAV, PCM 16-bit, mono Language Thai Source… See the full description on the dataset page: https://huggingface.co/datasets/Chalermdej/yodas2_sidon_th_tts.audiotext-to-speech100K<n<1M3 likes425 downloads4mo agoHugging Face25TTS-AGI /advanced-soundscapes-stage-1 Advanced Soundscapes Stage 1 — Raw Components (5M) This dataset contains Stage 1 output from the LAION Universal Audio Annotation Pipeline (UAAP) data generation plan. Contents 5,000 shards containing 5,000,000 soundscape recipes with raw audio components Each soundscape row includes: recipe.json — full recipe with timeline, events, loudness, speaker IDs, overlap/density settings spkN.flac / spkN.json — raw speech components + full source metadata musicN.flac /… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/advanced-soundscapes-stage-1.audioaudio-classification1M<n<10M0 likes424 downloads3mo agoHugging Face26ghananlpcommunity /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes410 downloads2mo agoHugging Face27serdarcaglar /turkish-tts-audiobooksgated Turkish TTS Audiobooks Turkish read-speech corpus for text-to-speech training, built from Turkish audiobook and spoken-article recordings by an automatic pipeline: VAD segmentation → technical QC → acoustic event tagging → DNSMOS → speaker embedding/consistency → double-pass Whisper ASR → text policy → leakage-free splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards. The pipeline that produced it — every stage, every threshold, the export and audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.audiotext-to-speech100K<n<1M9 likes378 downloads1mo agoHugging Face28thennal /indic_tts_ml Indic TTS Malayalam Speech Corpus The Malayalam subset of Indic TTS Corpus, taken from this Kaggle database. The corpus contains one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given in the repository. audiotext-to-speech1K<n<10K6 likes377 downloads4y agoHugging Face29LanguaMan /vieneu-tts-140h-dataset pnnbao-ump/VieNeu-TTS-140h Mô tả Dataset Dataset tiếng Việt chất lượng cao cho Text-to-Speech (TTS) với 74,858 mẫu audio và transcript được phonemize. Mục tiêu của mình là tạo bộ dataset chuẩn mực để finetune các model TTS hiện nay với chất lượng cao nhất. Mình thu thập audio chất lượng cao từ youtube, làm sạch nền, loại bỏ noise, dùng whisper-large-v3 để tạo transcription, sau đó cho Agent sửa lỗi chính tả và feedback lại cho con người. Bộ dữ liệu cũng được phonemize hóa… See the full description on the dataset page: https://huggingface.co/datasets/LanguaMan/vieneu-tts-140h-dataset.audiotext-to-speech10K<n<100K2 likes341 downloads5mo agoHugging Face30GianDiego /latam-spanish-speech-orpheus-tts-24khz LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready) Dataset Description This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate. The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.audiotext-to-speech10K<n<100K16 likes333 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.