datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cml-tts
Dataset Card for CML-TTS
Dataset Summary
CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG).
CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.voicehub-arena-seed-tts-eval
VoiceHub Arena — native TTS evaluations
Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA
speaker SIM and UTMOS22 measurements. The full campaign is still running.
Each generation method is evaluated separately using its publisher's native API.
Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target
diagnostic pilots are stored separately and must not be treated as full scores.
Interactive demo ·
Source code
Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.imagesseed-tts-eval-arrowuyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.seed-tts-eval
seed-tts-eval
A preprocessed copy of the seed-tts-eval test set, used by SGLang Omni for TTS benchmarking (WER and speed evaluation).
We thank the researchers of ByteDance for releasing the original evaluation data and methodology. This dataset simply reorganizes their test sets into a single Hugging Face repository for convenience.
Evaluation Sets
This dataset contains 5 evaluation sets across English and Chinese:
#
File
Language
Samples
Columns
Difficulty… See the full description on the dataset page: https://huggingface.co/datasets/zhaochenyang20/seed-tts-eval.ramanv-tts-all-raw
ramanv-tts-all-raw
Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages.
libritts_r_filtered
Dataset Card for Filtered LibriTTS-R
This is a filtered version of LibriTTS-R. It has been filtered based on two sources:
LibriTTS-R paper [1], which lists samples for which speech restoration have failed
LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected.
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately
585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.tts-de1qwen3-tts-polish-traininglistening_test
Listening Test Results for TTSDS2
This dataset contains all 11,000+ ratings collected for 20 synthetic speech systems for the TTSDS2 study (link coming soon).
The scores are MOS (Mean Opinion Score), CMOS (Comparative Mean Opinion Score) and SMOS (Speaker Similarity Mean Opinion Score).
All annotators included passed three attention checks throughout the survey.
mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
voicehub-arena-seed-tts-eval
VoiceHub Arena — full English Seed-TTS-Eval
35,904 synthesized WAV files: 33 model families × the same 1,088 target texts.
The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB.
All 198 shards and every WAV SHA256 were verified after backup.
Interactive leaderboard and all audio samples
· Source repository (access required).
Contents
audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.hifi-tts
Dataset Card for HiFiTTS
Hi-Fi Multi-Speaker English TTS Dataset (Hi-Fi TTS) is based on LibriVox's public domain audio books and Gutenberg Project texts.
open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.Bagpiper_TTS_SFT_Data
Bagpiper-TTS SFT Data
Release status: the validated Parquet release is being uploaded. The
homepage and metadata may appear before every large shard is committed.
Bagpiper-TTS SFT Data supports
Bagpiper-TTS, a universal
speech-synthesis model that interprets free-form natural-language requests,
plans the requested delivery, produces a rich textual caption, and synthesizes
the target audio.
The release is organized into the six applications used by the paper:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_TTS_SFT_Data.ISCSLP2026-CoT-TTS
ISCSLP 2026 CoT-TTS Dataset
Dataset Overview
This dataset is prepared for the ISCSLP 2026 CoT-TTS Challenge and is designed to support research on context-aware, expressive, and CoT-guided speech generation. It is constructed from speech-rich media sources, including films, TV dramas, radio dramas, and short dramas, where dialogue often contains rich conversational context, speaker interactions, scene changes, and emotional variation. Each sample is organized… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/ISCSLP2026-CoT-TTS.tts_leaderboard_screenshotslahgtna-libyan-ttsmajestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.indic-tts-966h
Indic-TTS-966h
Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV
clips with sentence-level transcripts in native scripts (natural English code-switching
preserved).
Subset
Clips
Hours
bengali
18,343
94.9
malayalam
30,548
192.5
marathi
34,327
213.4
punjabi
28,083
161.8
tamil
26,817
171.1
telugu
21,923
132.8
Columns: audio (24 kHz mono), file_name, transcript. One config per language:
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.tts_farm
Multilingual TTS/ASR Aggregated Dataset
Cleaned, deduplicated and loudness-normalized Arabic, Japanese, Korean, Turkish, and Vietnamese speech. The training columns are audio (16-bit PCM WAV, 22050 Hz) and text; the remaining columns contain quality and provenance metadata.
Mana-TTS
ManaTTS-Persian-Speech-Dataset
ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality audio (sampled at 44.1 kHz). Released under the permissive CC-0 license, this dataset is freely usable for both educational and commercial purposes.
Collected from Nasl-e-Mana magazine, the dataset covers a diverse range of topics, making it ideal for training robust text-to-speech (TTS) models. The release includes a fully transparent… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/Mana-TTS.tts-dataset-combinedTTS_ArenaTTS Arena's DB is SQLlite DB file. The above is just a summary query that should be useful for TTS developers to evaluate faults of their model.
Why no audio samples?
Unsafe. Cannot constantly oversee the output of uncontrolled HuggingFace Spaces. While it could be safeguarded by using an ASR model before uploading, something unwanted may still slip through.
Useful queries for TTS developers and evaluators
All votes mentioning specified TTS model:… See the full description on the dataset page: https://huggingface.co/datasets/Pendrokar/TTS_Arena.seed-tts-eval-minitts-datagen
GPT-OSS 120B native reasoning traces for TTS Datagen
Summary
This dataset contains 2,865 synthetic competitive-programming questions,
45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50
verified test cases per question (143,250 test cases total). Each solution
preserves the model's native reasoning trace separately from its final answer.
The reasoning was returned by MetaGen's native Dialog Completion interface as
dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.dahih-tts2-demucs-cleanedovos-tts-bench-massive-prompts
OVOS tts bench — massive-prompts
Synthesised clips (one per prompt) predictions of the registered
OVOS Plugin Arena
tts fighters over
OpenVoiceOS/massive-templates.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-tts-bench-massive-prompts.
