datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cml-tts
Dataset Card for CML-TTS
Dataset Summary
CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG).
CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.voicehub-arena-seed-tts-eval
VoiceHub Arena — native TTS evaluations
Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA
speaker SIM and UTMOS22 measurements. The full campaign is still running.
Each generation method is evaluated separately using its publisher's native API.
Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target
diagnostic pilots are stored separately and must not be treated as full scores.
Interactive demo ·
Source code
Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.uyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.seed-tts-eval
seed-tts-eval
A preprocessed copy of the seed-tts-eval test set, used by SGLang Omni for TTS benchmarking (WER and speed evaluation).
We thank the researchers of ByteDance for releasing the original evaluation data and methodology. This dataset simply reorganizes their test sets into a single Hugging Face repository for convenience.
Evaluation Sets
This dataset contains 5 evaluation sets across English and Chinese:
#
File
Language
Samples
Columns
Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/zhaochenyang20/seed-tts-eval.libritts_r_filtered
Dataset Card for Filtered LibriTTS-R
This is a filtered version of LibriTTS-R. It has been filtered based on two sources:
LibriTTS-R paper [1], which lists samples for which speech restoration have failed
LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected.
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately
585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.tts-de1listening_test
Listening Test Results for TTSDS2
This dataset contains all 11,000+ ratings collected for 20 synthetic speech systems for the TTSDS2 study (link coming soon).
The scores are MOS (Mean Opinion Score), CMOS (Comparative Mean Opinion Score) and SMOS (Speaker Similarity Mean Opinion Score).
All annotators included passed three attention checks throughout the survey.
mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
voicehub-arena-seed-tts-eval
VoiceHub Arena — full English Seed-TTS-Eval
35,904 synthesized WAV files: 33 model families × the same 1,088 target texts.
The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB.
All 198 shards and every WAV SHA256 were verified after backup.
Interactive leaderboard and all audio samples
· Source repository (access required).
Contents
audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.hifi-tts
Dataset Card for HiFiTTS
Hi-Fi Multi-Speaker English TTS Dataset (Hi-Fi TTS) is based on LibriVox's public domain audio books and Gutenberg Project texts.
Bagpiper_TTS_SFT_Data
Bagpiper-TTS SFT Data
Release status: the validated Parquet release is being uploaded. The
homepage and metadata may appear before every large shard is committed.
Bagpiper-TTS SFT Data supports
Bagpiper-TTS, a universal
speech-synthesis model that interprets free-form natural-language requests,
plans the requested delivery, produces a rich textual caption, and synthesizes
the target audio.
The release is organized into the six applications used by the paper:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_TTS_SFT_Data.lahgtna-libyan-ttsmajestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.indic-tts-966h
Indic-TTS-966h
Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV
clips with sentence-level transcripts in native scripts (natural English code-switching
preserved).
Subset
Clips
Hours
bengali
18,343
94.9
malayalam
30,548
192.5
marathi
34,327
213.4
punjabi
28,083
161.8
tamil
26,817
171.1
telugu
21,923
132.8
Columns: audio (24 kHz mono), file_name, transcript. One config per language:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.tts_farm
Multilingual TTS/ASR Aggregated Dataset
Cleaned, deduplicated and loudness-normalized Arabic, Japanese, Korean, Turkish, and Vietnamese speech. The training columns are audio (16-bit PCM WAV, 22050 Hz) and text; the remaining columns contain quality and provenance metadata.
seed-tts-eval-minicml-tts-filtered
Dataset Card for Filtred and CML-TTS
This dataset is a filtred version of a CML-TTS [1].
CML-TTS [1] CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in… See the full description on the dataset page: https://huggingface.co/datasets/PHBJT/cml-tts-filtered.dahih-tts2-demucs-cleanedagent-sft-stitch-zh-tts
agent-sft-stitch-zh-tts
Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted.
Configs
records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.cml-ttsTTS-Clean44k
TTS-Clean44k
A multilingual pool of verified-clean, wideband speech for training and evaluating speech
restoration / text-to-speech (TTS) models. Every utterance is independently checked on two
axes and stored as parquet with its per-utterance quality scores attached:
Native sample rate ≥ 44.1 kHz — measured per file with ffprobe, never trusting the
source's advertised rate. Anything below 44.1 kHz is dropped.
DNSMOS P.835 bak ≥ 3.644 — the background-noise MOS from the DNSMOS… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/TTS-Clean44k.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.ghana-english-tts-clean2
Ghana English TTS Filtered Clean v2
Filtered subset of ghananlpcommunity/ghana-english-tts-filtered using PANNs CNN14.
Filtering
Second-pass filtering with PANNs CNN14 (soundclassifier with music, applause, and speech tags):
Keep if: music_prob ≤ 0.2 AND applause_prob ≤ 0.2 AND speech_prob ≥ 0.5
Batch size 32 on NVIDIA H200, float16 inference
282,096 / 303,204 kept (93.0%)
Fields
corrected_text: utterance text
bytes: raw WAV bytes (16-bit PCM… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-english-tts-clean2.Malaysian-TTS
TTS
Malaysian Synthetic TTS dataset.
Generate using each Malaysian-F5-TTS-v2.
Each generation verified using esammahdi/ctc-forced-aligner.
Post-filter pitch using interactiveaudiolab/penn.
Speaker
Husein, 300 hours.
Shafiqah Idayu, 292 hours.
Anwar Ibrahim, 269 hours.
KP RTM Suhaimi Malay, 306 hours.
KP RTM Suhaimi Chinese, 192 hours.
Clean version
We trimmed start and end silents, and compressed at processed
Dataset uploaded as HuggingFace datasets… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-TTS.Env-TTS-Clean
Env-TTS-Clean
Environment-aware text-to-speech training corpus (clean release). Each row
pairs four short 24 kHz mono FLAC clips with aligned transcripts:
an environment sample (different speaker, same acoustic scene),
a speaker reference (same speaker as the target utterance),
a speaker-enhanced copy of the reference (MossFormer2 enhancement — or, for
the DDS source, the real clean-studio recording of the speaker reference),
the target speech to synthesise,
so a model can… See the full description on the dataset page: https://huggingface.co/datasets/humanify/Env-TTS-Clean.seed-tts-eval-50-arrowmls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.Synthetic-User-Turn-TTS
Synthetic Malaysian Telco Call-Centre Speech
Synthetic Malaysian call-centre customer utterances, as text and as speech.
The text is fully synthetic dialogue styled after real Malaysian ISP/telco ("Unifi")
call-centre recordings, containing no real customer data. The audio subsets take customer
(user) turns and voice them with a voice-conversion model, keeping only clips an ASR
round-trip confirms are accurate.
Subsets
subset
rows
content
default
4,260… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Synthetic-User-Turn-TTS.
