datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reazon-speech-v2-clone
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.tts-pretrain-clones-3m-mos
TTS Pretrain Clones (3M) — with DNSMOS
This is SynDataLab/tts-pretrain-clones-3m
with an added per-utterance dnsmos column (DNSMOS P.835 OVRL score, float32),
computed with the sig_bak_ovr.onnx model.
2,967,779 clone utterances across 2971 English speakers.
Sample rate: 44.1 kHz, WAV in Parquet
dnsmos: overall MOS quality estimate per utterance (higher is better)
Generated by echo-tts synthesizing English text on speaker latents
derived from Qwen3-TTS VoiceDesign base speakers.… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/tts-pretrain-clones-3m-mos.irodori-clones-3m-v2-no-emoji
Irodori TTS Clones v2 (3.29M)
3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2,
using the 10,000 reference voices from SynData-2/irodori-refs-10k-v2.
329 clones per ref voice, each with a unique Japanese conversational text.
Companion refs: SynData-2/irodori-refs-10k-v2.
Note: Bu dataset irodori-clones-3m-v2'nin emoji-temizlenmis kopyasidir. Audio bytes binary-identical; yalnizca text kolonundaki emojiler kaldirilmistir (emoji kutuphanesi, Japonca/CJK… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA/irodori-clones-3m-v2-no-emoji.QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-EmotionITA-Corpus Emotion Dataset (100 Japanese Female Voices)
彼のあだ名は言い得て妙だよね
11:A lower-pitched female voice with a strong core
ヒューズが飛んだ
100:A slightly quirky female voice that leaves a strong impression
Overview
This dataset contains 100 female voices generated with Qwen3-TTS.
Format: 24kHz mono WAV
Source: Link to designed voices
About ITA-Corpus Emotion
The text is based on the ITA-Corpus Emotion, a public domain dataset containing 100… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-Emotion.qwen_real_clonevox-cloned-data
CommonVoice Clones
This dataset consists of recordings taken from the CommonVoice english dataset.
Each voice and transcript are used as input to a voice cloner, and generate a cloned version of the voice and text.
TTS Models
We use the following high-scoring models from the TTS leaderboard:
playHT
metavoice
StyleTTSv2
XttsV2
Model Comparisons
To facilitate data exploration, check out this HF space 🤗, which allows you to listen to all clones from a given… See the full description on the dataset page: https://huggingface.co/datasets/jerpint/vox-cloned-data.Afrivoice_Kinyarwanda_ASR_clonecloneqwen-clones-4m-en4M conversational English speech clips generated with Qwen3-TTS-12Hz-1.7B-Base.
Reference speakers: SynData-2/qwen-ref-speakers-4k-en
neutral-batch-voice-cloneecho-clones-4m-en
echo-clones-4m-en
~4 M English TTS clone utterances generated with
EchoTTS (jordand/echo-tts-base).
Sample rate: 44 100 Hz, 16-bit PCM WAV stored in Parquet
Speakers: 4 000 reference speakers (spk_0000-spk_3999)
Text bucketing: quip (<=100 chars), mid (100-300), ramble (300-420)
Speaker assignment: round-robin -- text[i] -> spk_{i % 4000}
Companion datasets
Reference speakers: SynData-2/echo-ref-speakers-4k-en -- the 4 000 reference WAVs used as speaker… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/echo-clones-4m-en.cosyvoice-clone
LALM Emotional Vulnerability Dataset
Overview
This dataset contains synthesized malicious speech instructions across multiple emotions and intensity levels to evaluate the safety responsiveness of Large Audio-Language Models (LALMs). The dataset aims to examine how speaker emotion and intensity influence the safety and robustness of AI responses.
Dataset Composition
Total samples: 8,320
Emotion categories:
Neutral: 520 samples
Angry: 1560 samples… See the full description on the dataset page: https://huggingface.co/datasets/LALM-emotional-vulnerability/cosyvoice-clone.clonecosyvoice-cloneelise-clone
Custom Elise-like TTS Dataset
Converted on 2025-06-18T05:56:51Z.
Samples: 1043
Format : 10-s clips with text transcription (like MrDragonFox/Elise)
Structure
column
type
description
audio
audio
24kHz mono wav clip
text
string
transcription
NaiLong-Voice-Clone
奶龙语音克隆数据集
完整项目与 Demo 效果可参见 GitHub
如果这个数据集对你有帮助,欢迎在 GitHub 上点个 Star ⭐ 支持一下!
数据集介绍
数据集按处理阶段分为以下四部分:
1. raw_audio (原始采样)
处理方式:使用 Audacity 直接对视频素材进行录音,格式为 44.1kHz, 16-bit, Stereo。
说明:包含背景音、特效及多角色对话的非结构化原片素材,是整个流水线的起点。
2. vocal_only (人声分离)
处理方式:从 raw_audio 中使用 UVR5 的 MDX-Net 模型剥离背景音乐与噪音。
说明:利用 MDX-Net 模型提取出干净的人声轨道,为后续切片提供高信噪比素材。
3. sliced_vocal (自动化切片)
处理方式:基于停顿检测、音色突变及总时长控制,将 vocal_only 自动化切分为一系列短音频。… See the full description on the dataset page: https://huggingface.co/datasets/pengyichen/NaiLong-Voice-Clone.tts-zh-clone-bench-4model
tts-zh-clone-bench-4model
Chinese (Mandarin) TTS voice-cloning comparison across 4 models. Each row is a
reference->clone pair; the same index uses the same text across all models
(rows are sorted by index, so the 4 models for one text are adjacent — easy A/B).
Columns: index, model, ref_text, ref_audio, ref_dnsmos, clone_text, clone_audio, clone_dnsmos, clone_asr_cer. Audio is 24 kHz; *_dnsmos is the DNSMOS P.835 OVRL
score (higher is better, ~1-5); clone_asr_cer is the… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-zh-clone-bench-4model.tts-multiling-clone-6model
tts-multiling-clone-6model
Single voice -> EN/JA/ZH across 6 open-source TTS models (20 VCTK reference speakers, cross-lingual zero-shot clone).
Model
DNSMOS EN
DNSMOS JA
DNSMOS ZH
spk_sim EN
spk_sim JA
spk_sim ZH
cosyvoice
3.31
3.22
3.18
0.59
0.51
0.53
moss
3.18
3.16
3.19
0.71
0.52
0.53
omnivoice
3.21
3.28
3.23
0.74
0.62
0.65
qwen3
3.25
3.30
3.20
0.72
0.62
0.59
voxcpm
3.09
3.06
3.17
0.64
0.48
0.43
zonos
3.21
3.39
3.37
0.52
0.35
0.36
tts-pretrain-clones-3m
TTS Pretrain Clones (3M)
2,967,779 clone utterances across 2971 English speakers.
Sample rate: 44.1 kHz, WAV in Parquet
Generated by echo-tts synthesizing English text on speaker latents
derived from Qwen3-TTS VoiceDesign base speakers. Per speaker:
10 voice-clone latents × 100 texts. The first utterance of each
speaker (row 0) is published separately in the companion refs set.
Coverage: speakers 1-60 + 61 (partial, 749 rows) + 91-3000. Thirty
speakers (61's tail + 62-90) are… See the full description on the dataset page: https://huggingface.co/datasets/humair-experiments/tts-pretrain-clones-3m.cloneclonedVSrealyt3_chunked_speech_restorised_tts_train_clone_pairs
yt3_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_tts_train_clone_pairs.yt2_chunked_speech_restorised_tts_train_clone_pairs
yt2_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_tts_train_clone_pairs.plt-clone-datasetdefault_voices_chunked_speech_restorised_tts_train_clone_pairs
default_voices_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Uzbek TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: uz (Uzbek)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_speech_restorised_tts_train_clone_pairs.tbp_chunked_speech_restorised_tts_train_clone_pairs
tbp_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_speech_restorised_tts_train_clone_pairs.audiobook_chunked_speech_restorised_tts_train_clone_pairs
audiobook_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Uzbek TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: uz (Uzbek)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_speech_restorised_tts_train_clone_pairs.espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs
espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_tts_train_clone_pairs.yt4_chunked_speech_restorised_tts_train_clone_pairs
yt4_chunked_speech_restorised_tts_train_clone_pairs
This is a gated Russian TTS training clone-pair dataset.
It contains tokenized speaker-reference and target pairs for text-to-speech voice adaptation workflows.
Language
Primary language: ru (Russian)
Contents
audios/shard-*.tar: tokenized audio shards
txts/shard-*.jsonl: per-example metadata and text fields
data.lst: repository-relative shard manifest
tokenized_dataset_summary.json: upload-time… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_tts_train_clone_pairs.clone
