datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mana-TTS
ManaTTS-Persian-Speech-Dataset
ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality audio (sampled at 44.1 kHz). Released under the permissive CC-0 license, this dataset is freely usable for both educational and commercial purposes.
Collected from Nasl-e-Mana magazine, the dataset covers a diverse range of topics, making it ideal for training robust text-to-speech (TTS) models. The release includes a fully transparent… See the full description on the dataset page: https://huggingface.co/datasets/MahtaFetrat/Mana-TTS.tts-datagen
GPT-OSS 120B native reasoning traces for TTS Datagen
Summary
This dataset contains 2,865 synthetic competitive-programming questions,
45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50
verified test cases per question (143,250 test cases total). Each solution
preserves the model's native reasoning trace separately from its final answer.
The reasoning was returned by MetaGen's native Dialog Completion interface as
dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.Malaysian-TTS-v2
Malaysian TTS v2
Generate Malay and localize English for TTS dataset, currently only support 2 speakers, husein and idayu, where total audio is 4642.77 hours.
How to prepare the dataset
huggingface-cli download \
mesolitica/Malaysian-TTS-v2 \
--include "all-*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/STT-Normalizer \
--include "*husein*.zip" \
--exclude "*force*" \
--repo-type "dataset" \
--local-dir './'… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-TTS-v2.cml-tts-100h-cappedparler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.tts-realspeech-sft-en-de
LAION TTS Real-Speech SFT — English + German, emotion-balanced
1,947,272 real recorded utterances — no synthetic voices — selected from freely-licensed corpora
and balanced across 40 emotions x 2 languages. 6,996 hours,
313,844,544 MOSS frames (3,766,134,528 audio tokens), 79,337,527 aligned words.
Each row is a self-contained TTS example: a corrected procedural caption, the transcript with
word-level timestamps, the original audio, and the target MOSS-Audio-Tokenizer-v2 codes.… See the full description on the dataset page: https://huggingface.co/datasets/laion/tts-realspeech-sft-en-de.tts-realspeech-dpo-en-de
LAION TTS Real-Speech DPO — English + German
3,959,192 preference pairs built from laion/tts-realspeech-sft-en-de
(EN 1,974,660 · DE 1,984,532). Every chosen is a real recorded
utterance; every rejected is that same utterance corrupted in one of the two ways a
caption-conditioned TTS model actually fails: stopping early or running on.
from datasets import load_dataset
import numpy as np
ds = load_dataset("laion/tts-realspeech-dpo-en-de", split="train", streaming=True)
r =… See the full description on the dataset page: https://huggingface.co/datasets/laion/tts-realspeech-dpo-en-de.natori-irodori-tts-dataset
natori-irodori-tts-dataset
High-quality single-speaker Japanese speech dataset prepared for Irodori-TTS LoRA training from multiple long-form さなちゃんねる videos featuring 名取さな.
Dataset Summary
This repository is a merged export of four curated subsets derived from public YouTube playlists on さなちゃんねる.
Current top-level merged export:
40,995 total utterances
33,398 training examples
7,597 validation examples
43.4277 hours of speech
speaker id: natori_sana
The merged… See the full description on the dataset page: https://huggingface.co/datasets/argo11/natori-irodori-tts-dataset.round2-oss-matched
round2-oss-matched — 第二轮 4 组实验数据(每组 10 节点,共 40)
代码:repo 分支 claude/round2-matched-compute(先 git fetch origin && git merge origin/claude/round2-matched-compute)。
目录:
exp0_20b/node00..04/pool.jsonl # 实验 0:shard-05 修复重跑(20B)
exp0_120b/node00..04/pool.jsonl # 实验 0:同上(120B)
exp1_20b/node00..09/{seeds,budgets,pool}.jsonl # 实验 1:20B token 对齐独立采样
exp2_120b/node00..09/{seeds,budgets,pool}.jsonl # 实验 2:120B 同上
exp3_120b/node00..09/{ck_nonsat/,nonsat_seeds,budgets,pool… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round2-oss-matched.snac_llm_parler_ttsagent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.mls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.Mana-TTS
ManaTTS-Persian-Speech-Dataset
ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality audio (sampled at 44.1 kHz). Released under the permissive CC-0 license, this dataset is freely usable for both educational and commercial purposes.
Collected from Nasl-e-Mana magazine, the dataset covers a diverse range of topics, making it ideal for training robust text-to-speech (TTS) models. The release includes a fully transparent… See the full description on the dataset page: https://huggingface.co/datasets/mainkilora/Mana-TTS.multilingual-tts-voice-dataset
Multilingual TTS Voice Dataset
Multilingual speech and structured voice-control data for text-to-speech research and training.
The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually.
Configurations
audio: utterances with embedded audio.
prompt_specs: structured text and delivery specifications.
clone_pairs: same-speaker reference and target pairs.
voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for six language variants, built
from Common Voice 17.0 by coverage-driven selection rather than random sampling.
Every example pairs a reference clip of one speaker with a target text that
speaker never read, so a system is asked to clone a voice and produce new
speech, which is what zero-shot TTS is actually for.
Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.libritts-r-filtered-speaker-descriptions
Dataset Card for Annotated LibriTTS-R
This dataset is an annotated version of a filtered LibriTTS-R [1].
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 960 hours of read English speech at 24kHz sampling rate, published in 2019.
In the text_description column, it provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts-r-filtered-speaker-descriptions.tts-scaling-ladder-de-en
TTS Scaling Ladder DE/EN — eight nested, balanced, openly licensed speech tiers
A nested ladder of eight speech datasets for scaling-law experiments on TTS, voice-conversion and audio
foundation models. Every tier is a strict subset of the next; every tier is 50/50 German/English, within
each language 50/50 real/synthetic, and inside each of those four cells stratified to be uniform over 40
EmoNet emotion categories (pipeline A) and uniform over 57 VoiceNet voice dimensions × 10… See the full description on the dataset page: https://huggingface.co/datasets/laion/tts-scaling-ladder-de-en.laion-tts-annotated-v1
LAION TTS Annotated v1
107,563,551 annotated speech utterances across six subsets — with the audio, the codec tokens
and the annotations, all joined by one key.
283,681 audio-hours. Per utterance: the transcript with word-level timings, 40 emotion
intensities, 57 VoiceNet voice-character dimensions, four audio-quality heads, vocal-burst
detections with timings, and a natural-language caption describing the voice and the
delivery — plus the audio itself, its… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1.sqp-tts-en
SQP TTS (English)
Synthesized speech for SQPsychConv_qwen-2.5, a synthetic CBT therapist-client
dialogue dataset (English).
Each configuration below corresponds to one TTS model. Load a single model
with:
from datasets import load_dataset
ds = load_dataset("sinselm/sqp-tts-en", "qwen3-tts")
Models included
qwen3-tts: https://huggingface.co/Qwen/Qwen3-TTS-12Hz-0.6B-Base
cosyvoice: https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512
fishaudio:… See the full description on the dataset page: https://huggingface.co/datasets/marleen-snsl/sqp-tts-en.round4-independent
ROUND 4 — combined r2+r3 pool, oracle-fb + independent, WITH REASONING RETENTION
Cut 2026-08-12. The first generation round whose outputs keep the gpt-oss
analysis channel (reasoning) on disk — see tts-sft/docs/REASONING_RETENTION.md.
Rounds 2–3 saved only the harmony final channel; their reasoning is unrecoverable.
What round 4 is
The combined pool — round-2 rerun pool (4,322 problems, apps-*/cc-*) +
round-3 pool (2,833 problems, cc3-*), zero id overlap, 7,155… See the full description on the dataset page: https://huggingface.co/datasets/tts-sft/round4-independent.Env-TTS-SD-Corpus-24K-Enhancedcml-tts-filtered-annotated
Dataset Card for Filtred and annotated CML TTS
This dataset is an annotated and filtred version of a CML-TTS [1].
CML-TTS [1] CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The… See the full description on the dataset page: https://huggingface.co/datasets/PHBJT/cml-tts-filtered-annotated.JA_Emilia_Yodas_ScribeEvents
JA Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/JA_Emilia_Yodas_266h containing only samples with ElevenLabs Scribe v1 audio events.
Changes from source
Filtered to rows where events_scribe is non-empty (4433 rows kept)
Bracket format unified: (event) in text_scribe replaced with [event]
Event types include
Vocal bursts: <laughs>, <sighs>, <clears throat>, etc.
Background: <background noise>, <music>, etc.
Other: <pause>, <unintelligible>… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/JA_Emilia_Yodas_ScribeEvents.mls-eng-10k-tags_tagged_10k_generated
Dataset Card for Annotations of 10K hours of English MLS
This dataset consists in annotations of a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-10k-tags_tagged_10k_generated.egyptian-arabic-tts-diacritized
Egyptian Arabic TTS Corpus (Diacritized)
97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with
diacritized transcripts — the short vowels that Arabic script does not
write.
Why diacritics
Arabic is an abjad: short vowels are unwritten, so كتب may be kataba,
kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which,
and guesses — which native listeners hear as a foreign accent with constant
mispronunciation.
This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.libritts_r_tags_tagged_10k_generated
Dataset Card for Annotated LibriTTS-R
This dataset is an annotated version of LibriTTS-R [1]. LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately 960 hours of read English speech at 24kHz sampling rate, published in 2019.
In the text_description column, it provides natural language annotations on the characteristics of speakers and utterances, that have been generated using the Data-Speech repository.… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_tags_tagged_10k_generated.ar-sa-tts-speakers-synthetic
Deprecated -- consolidated
This repo's data files stay in place, but the rows now live as named config(s) on the single Salesteq synthetic-speech dataset:
speakers-v1 on Salesteq/ar-sa-tts-corpus-synthetic
Load from there rather than this repo going forward.
ar-sa-tts-speakers — multi-speaker Najdi Arabic TTS (synthetic)
Ten single-speaker synthetic Najdi Arabic TTS sets unified into one dataset, distinguished by the speaker column. 91,350 clips across 10… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/ar-sa-tts-speakers-synthetic.sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.laion-tts-annotated-v1-research
LAION TTS Annotated v1 — research subsets
29,739,936 annotated speech utterances across three subsets — with the audio, the codec tokens
and the complete annotation stack.
The audio in this repository comes from podcasts that are openly available on the internet and
consists of short snippets only. We cannot redistribute the audio itself, so it is made available
here for non-commercial research use by collaboration partners within our TTS research.
The other six subsets of this… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-tts-annotated-v1-research.
