datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uyghur-common-voice-tts
Uyghur Common Voice TTS Dataset
A cleaned and processed Text-to-Speech (TTS) dataset for the Uyghur language, derived from Mozilla Common Voice.
Dataset Summary
Property
Value
Language
Uyghur (ug)
Total Samples
43,054
Train Samples
40,901
Validation Samples
2,153
Audio Format
WAV
Source
Mozilla Common Voice
License
CC0-1.0
Dataset Structure
/
├── train.jsonl # Training data (40,901 samples)
├── val.jsonl #… See the full description on the dataset page: https://huggingface.co/datasets/anke01/uyghur-common-voice-tts.dahih-tts2-demucs-cleanedraw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16kraw_tts_esc_ESPnet_espnet_mls-multi_soundstream_16krequestsround5-sv-cells
round5-sv-cells — SELF-VERIFIER feedback cells (sv7 / sv7d)
2026-08-16. Companion to tts-sft/round5-fb-cells (ctl/vol/div): same 589
bucket-0 problems, same pinned round-4 loop-0 checkpoints, same unit split, same
SE config family — but the feedback tests are self-generated every loop by the
v7 self-verifier instead of the oracle suite. The oracle cells are the
controls; together they measure, at scale, how much of feedback-SE's bucket-0
reach and densification survives when the… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round5-sv-cells.round2-oss-matched
round2-oss-matched — 第二轮 4 组实验数据(每组 10 节点,共 40)
代码:repo 分支 claude/round2-matched-compute(先 git fetch origin && git merge origin/claude/round2-matched-compute)。
目录:
exp0_20b/node00..04/pool.jsonl # 实验 0:shard-05 修复重跑(20B)
exp0_120b/node00..04/pool.jsonl # 实验 0:同上(120B)
exp1_20b/node00..09/{seeds,budgets,pool}.jsonl # 实验 1:20B token 对齐独立采样
exp2_120b/node00..09/{seeds,budgets,pool}.jsonl # 实验 2:120B 同上
exp3_120b/node00..09/{ck_nonsat/,nonsat_seeds,budgets,pool… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round2-oss-matched.raw_tts_esc_ESPnet_espnet_mls-english_soundstream_16kTTS_eval_datasets
TTS evaluation datasets
This repository contains three testsets for zero-shot TTS models:
dialog_testset: Chinese and English testsets for spoken dialogue generation models, introduced in paper ZipVoice-Dialog.
librispeech_pc_testset: English testset for zero-shot TTS models, introduced in paper F5-TTS.
seedtts_testset: Chinese and English testsets for zero-shot TTS models, introduced in paper Seed-TTS.
minimax_multilingual_24: 24-language testset for zero-shot TTS models… See the full description on the dataset page: https://huggingface.co/datasets/k2-fsa/TTS_eval_datasets.TTS-Voice-Design-Benchmark
TTS Voice Design Benchmark
🏆 Leaderboard | 🛠️ Evaluation Suite
TTS Voice Design is a high-quality benchmark of 1,000 character voice-design
tasks spanning a broad range of media genres and real-world creative use
cases. It evaluates whether a text-to-speech model can turn an open-ended
character profile into a distinctive, appropriate, and usable voice.
Unlike benchmarks built around a fixed set of speakers or isolated acoustic
attributes, this dataset covers complete… See the full description on the dataset page: https://huggingface.co/datasets/BreezeBlue/TTS-Voice-Design-Benchmark.round4-independent
ROUND 4 — combined r2+r3 pool, oracle-fb + independent, WITH REASONING RETENTION
Cut 2026-08-12. The first generation round whose outputs keep the gpt-oss
analysis channel (reasoning) on disk — see tts-sft/docs/REASONING_RETENTION.md.
Rounds 2–3 saved only the harmony final channel; their reasoning is unrecoverable.
What round 4 is
The combined pool — round-2 rerun pool (4,322 problems, apps-*/cc-*) +
round-3 pool (2,833 problems, cc3-*), zero id overlap, 7,155… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round4-independent.arnum-tts
ArNum-TTS — does the number survive?
Author: Syamjith NK
License: CC BY 4.0 (data) · MIT (code)
A small evaluation set for one narrow, practical question: when an Arabic
text-to-speech system reads a sentence containing a number, can a listener recover
the number?
Not naturalness. Not accent. Not prosody. A voice can be beautiful and still be
unusable for anything containing a date, a price or a percentage — and that failure
is invisible in every TTS demo, because demos do not… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arnum-tts.Kyrg-TTSdarija-tts-8400
Darija TTS 8400
Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV.
All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings.
Write-up of how this data was used: Training a Voice.
At a glance
Clips / hours
8,400 / 20.73
Unique texts
4,800
Voice
Kore (1 speaker)
Sample rate
24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.round5-abcd-cells
Round-5 self-verifier D-cut — four new-verifier mechanisms (2026-08-18)
Four opt-in verifier mechanisms on top of the v7 stack (official-sample anchoring +
validity probes + wb-cands 8 + wb-certify), one arm each, on the 295-problem
mechanism-screening slice (u00+u01 of the 589 pool; same problems, budgets,
loop-0 population and GENSEED as the observed sv cells — rows are directly
comparable to the sv7/sv7d/pw7/sel7/sum7 screen table).
arm
flag
mechanism
bru7… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round5-abcd-cells.rgad-crosslingual-tts-10h
RGAD Cross-Lingual TTS 10h
This is a 10-hour cross-lingual TTS dataset for prompt-conditioned Chinese TTS fine-tuning.
Format
The dataset contains:
train.jsonl
dev.jsonl
metadata.csv
audio/prompts/*.wav
audio/targets/*.wav
Each JSONL row has this format:
{"id":"sample_000001","prompt_wav":"audio/prompts/sample_000001.wav","target_wav":"audio/targets/sample_000001.wav","text":"中文目标文本。","prompt_language":"en-US","target_language":"zh-CN"… See the full description on the dataset page: https://huggingface.co/datasets/isabeth/rgad-crosslingual-tts-10h.tts_llmsr_data
tts_llmsr_data
LDM-TTS-Base-SFT-19K
LDM-TTS-Base-SFT-19K
Supervised fine-tuning (SFT) corpus for Large Discovery Models (LDM): a dataset that
distils an acquisition-guided, test-time search policy into a language-model proposer so
that a single forward pass emulates a full model-based optimization loop.
Dataset Summary
An LDM couples three components in a recurrent generate → select → evaluate → update
loop: an LLM that proposes candidate experiments, a probabilistic surrogate that maps
observations… See the full description on the dataset page: https://huggingface.co/datasets/Yangtze-ailab/LDM-TTS-Base-SFT-19K.naijavoices_dataset_85_hours_tts_bestFull dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic metrics\color{Blue}{\large \textbf{Full dataset re-upload with more statistics as well as filtering scripts that give the top x files or best x hours for tts based and calculations from acoustic… See the full description on the dataset page: https://huggingface.co/datasets/David-A-Amoo/naijavoices_dataset_85_hours_tts_best.orpheus-tts-dataset-preserving-personalityMED-TTS
MED-TTS
MED-TTS (Multi-Emotion and Duration-annotated Text dataset for TTS) is a multilingual emotional transition dataset designed for expressive speech, audio, and multimodal generation research. Unlike conventional emotion datasets that assign a single static emotion label to an utterance, MED-TTS focuses on dynamic emotion flow within an utterance through segment-level annotations.
The dataset contains Chinese and English text samples annotated with:
utterance-level emotion… See the full description on the dataset page: https://huggingface.co/datasets/Chanson-0803/MED-TTS.atc-tts-mos-ratingsAzure-TTS-Yasmin-WikipediaAzure-TTS-Osman-WikipediaXijinping-TTS-Voicebank
习近平音源
所有声音资料来自公开影像,属于公有领域目前有 1h30m 的截取后声音,足够进行 Fine-tuning
Usage
按句截取
python -m pip install -r requirement.txt
python split.py
新增声音资料后,使用 Whisper 产生带有时间标记的 JSON 档,并手动复制到 ./voice/[FILE].json
export OPENAI_API_KEY="API_KEY_HERE"
python whisper.py ./[FILE].[AUDIO_EXTENSION]
产生 Bert-VITS2 微调所需的 esd.list 档案
python index_to_list.py
links_to_pocasts_lecture_and_shows_for_tts
This dataset was made by Charan from our Discord community. Thank you, very much. :)
license: apache-2.0
atc-tts-llm-mos-ratingsqwen3-tts-preset-voices
Qwen3-TTS preset voice embeddings
The 9 named speakers from Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice packaged as a small sidecar bundle usable with the -Base checkpoint.
bundle.safetensors — 9 × 2048-d bfloat16 rows, ~37 KB total
bundle.json — metadata (speaker name → spk_id, gender, supported languages)
Each row is lifted from talker.model.codec_embedding.weight in the CustomVoice checkpoint at the speaker-ID index from its config.json. With these rows, you can:
Deploy only… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-tts-preset-voices.interspeech2024_discrete_speech_tts_resultsyodas-tts
