CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ylacombe /cml-tts Dataset Card for CML-TTS Dataset Summary CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.audiotext-to-speech1M<n<10M36 likes118k downloads3y agoHugging Face02parler-tts /libritts_r_filtered Dataset Card for Filtered LibriTTS-R This is a filtered version of LibriTTS-R. It has been filtered based on two sources: LibriTTS-R paper [1], which lists samples for which speech restoration have failed LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected. LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.audiotext-to-speech100K<n<1M24 likes4.7k downloads2y agoHugging Face03lingamvamshikrishnareddy /ramanv-tts-all-rawgated ramanv-tts-all-raw Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages. textautomatic-speech-recognition1M<n<10M0 likes4.4k downloads12d agoHugging Face04ttsds /listening_test Listening Test Results for TTSDS2 This dataset contains all 11,000+ ratings collected for 20 synthetic speech systems for the TTSDS2 study (link coming soon). The scores are MOS (Mean Opinion Score), CMOS (Comparative Mean Opinion Score) and SMOS (Speaker Similarity Mean Opinion Score). All annotators included passed three attention checks throughout the survey. audioaudio-classification10K<n<100K3 likes3.9k downloads1y agoHugging Face05espnet /Bagpiper_TTS_SFT_Data Bagpiper-TTS SFT Data Release status: the validated Parquet release is being uploaded. The homepage and metadata may appear before every large shard is committed. Bagpiper-TTS SFT Data supports Bagpiper-TTS, a universal speech-synthesis model that interprets free-form natural-language requests, plans the requested delivery, produces a rich textual caption, and synthesizes the target audio. The release is organized into the six applications used by the paper:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_TTS_SFT_Data.audiotext-to-speech100K<n<1M0 likes3k downloads2mo agoHugging Face06multilingual-tts /open-bible OpenBibleTTS OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license. Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources Source: Open Bible (CC BY-SA) Languages Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.audiotext-to-speech1M<n<10M1 likes2.9k downloads3mo agoHugging Face07MikhailT /hifi-tts Dataset Card for HiFiTTS Hi-Fi Multi-Speaker English TTS Dataset (Hi-Fi TTS) is based on LibriVox's public domain audio books and Gutenberg Project texts. audiotext-to-speech100K<n<1M38 likes2.8k downloads3y agoHugging Face08psk /indic-tts-966h Indic-TTS-966h Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV clips with sentence-level transcripts in native scripts (natural English code-switching preserved). Subset Clips Hours bengali 18,343 94.9 malayalam 30,548 192.5 marathi 34,327 213.4 punjabi 28,083 161.8 tamil 26,817 171.1 telugu 21,923 132.8 Columns: audio (24 kHz mono), file_name, transcript. One config per language: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.audio100K<n<1M6 likes2.5k downloads2mo agoHugging Face09parler-tts /mls_eng Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.audioautomatic-speech-recognition10M<n<100M40 likes2.3k downloads2y agoHugging Face10MahtaFetrat /Mana-TTS ManaTTS-Persian-Speech-Dataset ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality audio (sampled at 44.1 kHz). Released under the permissive CC-0 license, this dataset is freely usable for both educational and commercial purposes. Collected from Nasl-e-Mana magazine, the dataset covers a diverse range of topics, making it ideal for training robust text-to-speech (TTS) models. The release includes a fully transparent… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/Mana-TTS.tabular10K<n<100K29 likes2.3k downloads1y agoHugging Face11RidheshBhati /tts_farm Multilingual TTS/ASR Aggregated Dataset Cleaned, deduplicated and loudness-normalized Arabic, Japanese, Korean, Turkish, and Vietnamese speech. The training columns are audio (16-bit PCM WAV, 22050 Hz) and text; the remaining columns contain quality and provenance metadata. audiotext-to-speech100K<n<1M2 likes2.2k downloads2mo agoHugging Face12PHBJT /cml-tts-filtered Dataset Card for Filtred and CML-TTS This dataset is a filtred version of a CML-TTS [1]. CML-TTS [1] CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in… See the full description on the dataset page: https://huggingface.co/datasets/PHBJT/cml-tts-filtered.audiotext-to-speech1M<n<10M4 likes1.7k downloads2y agoHugging Face13voidful /agent-sft-stitch-zh-tts agent-sft-stitch-zh-tts Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted. Configs records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.audiotext-to-speech100K<n<1M0 likes1.5k downloads3mo agoHugging Face14PHBJT /cml-ttsaudio1M<n<10M0 likes1.5k downloads2y agoHugging Face15ghanaopenai /new-twi-tts-aligned-ipa new-twi-tts-aligned + IPA phonemes ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme transcription for every clip, produced with ghananlpcommunity/ghana-speech-phoneme-asr. Audio included — this is self-contained, no join with the source dataset needed. Contents split clips hours phoneme units mean units/clip test 16,140 17.24 663,140 41.1 train 145,258 155.21 5,945,389 40.9 Columns column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.audioautomatic-speech-recognition100K<n<1M0 likes1.5k downloads2mo agoHugging Face16humanify /Env-TTS-Clean Env-TTS-Clean Environment-aware text-to-speech training corpus (clean release). Each row pairs four short 24 kHz mono FLAC clips with aligned transcripts: an environment sample (different speaker, same acoustic scene), a speaker reference (same speaker as the target utterance), a speaker-enhanced copy of the reference (MossFormer2 enhancement — or, for the DDS source, the real clean-studio recording of the speaker reference), the target speech to synthesise, so a model can… See the full description on the dataset page: https://huggingface.co/datasets/humanify/Env-TTS-Clean.audio100K<n<1M0 likes1.5k downloads2mo agoHugging Face17Scicom-intl /TTS-Clean44k TTS-Clean44k A multilingual pool of verified-clean, wideband speech for training and evaluating speech restoration / text-to-speech (TTS) models. Every utterance is independently checked on two axes and stored as parquet with its per-utterance quality scores attached: Native sample rate ≥ 44.1 kHz — measured per file with ffprobe, never trusting the source's advertised rate. Anything below 44.1 kHz is dropped. DNSMOS P.835 bak ≥ 3.644 — the background-noise MOS from the DNSMOS… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/TTS-Clean44k.audio1M<n<10M1 likes1.5k downloads2mo agoHugging Face18jigsaws-stomper /cml-tts-100h-cappedtabular100K<n<1M0 likes1.2k downloads6mo agoHugging Face19ghanaopenai /ghana-english-tts-clean2 Ghana English TTS Filtered Clean v2 Filtered subset of ghananlpcommunity/ghana-english-tts-filtered using PANNs CNN14. Filtering Second-pass filtering with PANNs CNN14 (soundclassifier with music, applause, and speech tags): Keep if: music_prob ≤ 0.2 AND applause_prob ≤ 0.2 AND speech_prob ≥ 0.5 Batch size 32 on NVIDIA H200, float16 inference 282,096 / 303,204 kept (93.0%) Fields corrected_text: utterance text bytes: raw WAV bytes (16-bit PCM… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-tts-clean2.audiotext-to-speech100K<n<1M0 likes1.2k downloads2mo agoHugging Face20parler-tts /mls_eng_10k Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.audioautomatic-speech-recognition1M<n<10M31 likes1.2k downloads2y agoHugging Face21mesolitica /Malaysian-TTS-v2 Malaysian TTS v2 Generate Malay and localize English for TTS dataset, currently only support 2 speakers, husein and idayu, where total audio is 4642.77 hours. How to prepare the dataset huggingface-cli download \ mesolitica/Malaysian-TTS-v2 \ --include "all-*.zip" \ --repo-type "dataset" \ --local-dir './' huggingface-cli download \ mesolitica/STT-Normalizer \ --include "*husein*.zip" \ --exclude "*force*" \ --repo-type "dataset" \ --local-dir './'… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-TTS-v2.tabular1M<n<10M2 likes1.2k downloads1y agoHugging Face22mesolitica /Malaysian-TTS TTS Malaysian Synthetic TTS dataset. Generate using each Malaysian-F5-TTS-v2. Each generation verified using esammahdi/ctc-forced-aligner. Post-filter pitch using interactiveaudiolab/penn. Speaker Husein, 300 hours. Shafiqah Idayu, 292 hours. Anwar Ibrahim, 269 hours. KP RTM Suhaimi Malay, 306 hours. KP RTM Suhaimi Chinese, 192 hours. Clean version We trimmed start and end silents, and compressed at processed Dataset uploaded as HuggingFace datasets… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-TTS.audio100K<n<1M0 likes1.2k downloads1y agoHugging Face23zhaochenyang20 /seed-tts-eval-50-arrowaudion<1K0 likes1.1k downloads4mo agoHugging Face24Scicom-intl /Synthetic-User-Turn-TTS Synthetic Malaysian Telco Call-Centre Speech Synthetic Malaysian call-centre customer utterances, as text and as speech. The text is fully synthetic dialogue styled after real Malaysian ISP/telco ("Unifi") call-centre recordings, containing no real customer data. The audio subsets take customer (user) turns and voice them with a voice-conversion model, keeping only clips an ASR round-trip confirms are accurate. Subsets subset rows content default 4,260… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Synthetic-User-Turn-TTS.audio100K<n<1M0 likes982 downloads21h agoHugging Face25datadriven-company /WolneLektury-TTS-Polish WolneLektury-TTS-Polish A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition. Data Source Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors. Dataset Statistics Metric Value Total samples 383,710 Total duration 997 hours Unique narrators 1207 Male samples 294,756 (767h) Female samples 88,945 (230h) Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.audiotext-to-speech100K<n<1M2 likes901 downloads8mo agoHugging Face26ghanaopenai /ewe-bible-audio-text-tts This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi 16-Word Speech Segments 48775 speech-text pairs split from long recordings. Processing pipeline Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.audioautomatic-speech-recognition10K<n<100K0 likes894 downloads3mo agoHugging Face27RidheshBhati /Indic-total-New-TTS-Merge Indic Total TTS Merge Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration. Languages assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu Columns audio: Audio data text: Transcript text duration: Duration in seconds (all >= 3.0s) language: Language name audio100K<n<1M1 likes885 downloads7mo agoHugging Face28alvanlii /cantonese-youtube-ttsgated Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet Filtered out: Overlapped voices, detected using pyannote/speaker-diarization-3.1 Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.audiotext-to-speech1M<n<10M3 likes873 downloads6mo agoHugging Face29ShiniChien /TTSDistil-Phonologygated Overview This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training. The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization. This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil-Phonology.audiotext-to-speech100K<n<1M1 likes858 downloads2mo agoHugging Face30PrakashPask /tts_indicaudio10K<n<100K1 likes835 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.