datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.Multilingual-TTS
Multilingual-TTS
A large multilingual corpus for pretraining TTS/STT models, gathered and normalized from 230+ public sources. ~191k hours of audio across 150+ languages, organized into 1,544 dataset configs and tokenized to 34.5B NeuCodec speech tokens (50 Hz, single-codebook) over 111.1M clips.
Each config is one source dataset normalized to rows of {audio_filename, text, speaker}:
audio_filename — clip path inside that config's <config>_audio.zip (mono MP3).
text —… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.Multilingual-TTS-language
Multilingual-TTS-language
malaysia-ai/Multilingual-TTS
with two extra columns:
column
description
audio_filename, text, speaker
unchanged from malaysia-ai/Multilingual-TTS
language
language detected from the text (transcript) column of every row
post-normalized
text after rule-based punctuation / capitalization normalization
All original columns and the file/folder layout are preserved: 1552 subsets /
1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.Normalized-Multilingual-TTS
Normalized Multilingual TTS
Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct.
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
qwen3-tts-multilingual-emotional-speechne-tts-coqui-multilingual
NE-TTS Coqui Multilingual
Multilingual TTS dataset for 15 North East Indian languages, formatted for Coqui-AI VITS multilingual training. Contains 61,943 clips / 83.5 hours at 22050Hz (SNR >= 20dB only).
Languages
ISO
Language
Clips
Hours
grt
Garo
24,772
29.6
ccp
Chakma
10,689
14.3
nag
Nagamese
9,688
14.5
lus
Mizo
8,554
14.5
nnp
Wancho
5,081
6.3
trp
Kokborok
1,237
1.7
clk
Idu Mishmi
602
0.7
mjw
Karbi
373
0.4
nre
Rengma
258
0.4
nri… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-coqui-multilingual.multipacking-multilingual-tts-10k-qwen3
multipacking-multilingual-tts-10k-qwen3
This is Mosaic format for https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS including multipacking max 10240 context length using Qwen3 tokenizer, so you can plug and play to train it.
For training example, you can check https://huggingface.co/malaysia-ai/Qwen3-1.7B-Multilingual-TTS
Multilingual-TTS-Voice-Conversion
Multilingual-TTS-Voice-Conversion
Convert TTS dataset to become Voice Conversion dataset by doing similarity combination. Multilingual TTS comes from malaysia-ai/Multilingual-TTS
Source code
Source code at https://github.com/Scicom-AI-Enterprise-Organization/Multilingual-TTS/tree/main/vc-tts
Acknowledgement
Special thanks to https://www.scitix.ai/ for H100 Node!
multilingual-tts-voice-dataset
Multilingual TTS Voice Dataset
Multilingual speech and structured voice-control data for text-to-speech research and training.
The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually.
Configurations
audio: utterances with embedded audio.
prompt_specs: structured text and delivery specifications.
clone_pairs: same-speaker reference and target pairs.
voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for six language variants, built
from Common Voice 17.0 by coverage-driven selection rather than random sampling.
Every example pairs a reference clip of one speaker with a target text that
speaker never read, so a system is asked to clone a voice and produce new
speech, which is what zero-shot TTS is actually for.
Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.melo-tts-ghana-multilingual
Melo-TTS multilingual Ghana training dataset
Ghanaian multilingual TTS training data for a MeloTTS fine-tune, phonemised with
ghananlpcommunity/xlsr-twi-codeswitch-ipa.
Languages
twi-asantee — Asante Twi (AfriSpeech open-bible-speech-african)
twi-akuapem — Akuapem Twi (AfriSpeech open-bible-speech-african)
ewe — Ewe (AfriSpeech open-bible-speech-african)
twi-en-codeswitch — Twi-English code-switched speech (ghananlpcommunity)
ghana-english —… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/melo-tts-ghana-multilingual.multilingual-tts
Before Anything and Everything ⚱
In the time of writing this Dataset Card, 17,490 18,412 civilian has been killed in Palestine (7,870 8,000 are children and 6,121 6,200 are women).
Seek any non-profit organization to help them with what you can (For myself, I use Mersal) 🇵🇸
Dataset Description
The Multilingual TTS dataset is an exceptional compilation of text-to-speech (TTS) samples, meticulously crafted to showcase the richness and diversity of human languages.… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/multilingual-tts.multilingual-synthetic-tts
Multilingual Synthetic TTS Dataset
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale synthetic multilingual speech dataset — 68,677 clips across
9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base
using zero-shot voice cloning from 5 reference speakers.
Intended for training and evaluating TTS, ASR, voice conversion, and
multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/YoonSeon/TTS-Multilingual-Test-Set.Multilingual-TTS-DNSMOS
Multilingual TTS — DNSMOS Filtered
A quality-filtered subset of malaysia-ai/Multilingual-TTS, retaining only audio samples that score OVRL ≥ 3.2 on the DNSMOS non-intrusive speech quality metric.
4,911,906 samples across multiple languages and TTS sources (84++ subset) remain after filtering.
Dataset Structure
Each record in the dataset contains the following fields:
Column
Type
Description
audio_filename
string
Relative path to the audio file within its… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-TTS-DNSMOS.mossv1.5-tts-multilingual-test
mossv1.5-tts-multilingual-test
5 reference voices (spk_1..spk_5) -> EN/IT/JA/TH/ZH zero-shot clone, MOSS-TTS-v1.5. One row per voice, per-language audio/text/metric columns.
Lang
DNSMOS
spk_sim
n
en
3.82
0.93
5
it
3.93
0.80
5
ja
3.80
0.85
5
th
3.86
0.81
5
zh
3.80
0.83
5
Reference audio: Expresso (CC BY-NC 4.0, non-commercial).
multilingual-tts
Before Anything and Everything ⚱
In the time of writing this Dataset Card, 17,490 18,412 civilian has been killed in Palestine (7,870 8,000 are children and 6,121 6,200 are women).
Seek any non-profit organization to help them with what you can (For myself, I use Mersal) 🇵🇸
Dataset Description
The Multilingual TTS dataset is an exceptional compilation of text-to-speech (TTS) samples, meticulously crafted to showcase the richness and diversity of human languages.… See the full description on the dataset page: https://huggingface.co/datasets/huayidu139/multilingual-tts.indic-multilingual-tts-v2
Indic Multilingual TTS v2
Combined dataset for training multilingual Indian-language TTS models, specifically prepared for Spark-TTS BiCodec and LLM fine-tuning.
Dataset Summary
Metric
Value
Total samples
556,524 (550K train + 6.5K val)
Total duration
~1,251 hours (1,237h train + 14h val)
Languages
13 (12 Indic + Indian English)
Sources
IndicVoices-R, Rasa, IndicTTS
Audio format
WAV, 16kHz mono
Emotion labels
6 emotions + neutral + domain tags… See the full description on the dataset page: https://huggingface.co/datasets/kapilkarda/indic-multilingual-tts-v2.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/jeshica/TTS-Multilingual-Test-Set.Malian-multilingual-SNAC-TTS-datasetmultilingual-synthetic-tts-de
multilingual-synthetic-tts-de
Mirror of the exact bmcore v24 holdout subset. Source: Reubencf/multilingual-synthetic-tts.
Local count agrees with published language count. Source card states reference-speaker consent and research/noncommercial use only; preserve restriction.
Contains 8998 synthetic audio files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: 332a32ddb6ff09935516cf5bf24a7fb0a5ec066c.
Original source… See the full description on the dataset page: https://huggingface.co/datasets/34data/multilingual-synthetic-tts-de.TTS_Multilingual_Data
Documentation Dataset: TTS_Multilingual_Data
Dataset Summary
This large-scale multilingual corpus is designed for linguistic analysis and the development of speech processing models. It supports tasks such as Text-to-Speech (TTS), Automatic Speech Recognition (ASR), and speaker identification. Structured in Parquet format, it serves as a key resource for training and evaluating models, using metrics tailored to ASR and speech technologies.
Thematic Categories… See the full description on the dataset page: https://huggingface.co/datasets/Databoost/TTS_Multilingual_Data.multilingual-tts-corpus
Multilingual TTS Corpus
A multilingual text-to-speech dataset containing audio recordings with text transcriptions across multiple languages. This is the first batch; more languages will be added over time.
Languages (Batch 1)
Subset
Language
Audio Files
Annotation Format
Source
ru-tts/
Russian
10
JSONL (single file)
俄语TTS(文本标注及语音)基石数据集 #294
ru-speech/
Russian
5
JSONL (single file)
俄语高质量语音音频语料库 #481
vi-speech/
Vietnamese
5
JSONL (single file)
越南语高质量语音音频语料库… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multilingual-tts-corpus.Malian-multilingual-SNAC-TTS-dataset-maya1multilingual-synthetic-tts-en
multilingual-synthetic-tts-en
Mirror of the exact bmcore v24 holdout subset. Source: Reubencf/multilingual-synthetic-tts.
Local count agrees with published language count. Source card states reference-speaker consent and research/noncommercial use only; preserve restriction.
Contains 5000 synthetic audio files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: 332a32ddb6ff09935516cf5bf24a7fb0a5ec066c.
Original source… See the full description on the dataset page: https://huggingface.co/datasets/34data/multilingual-synthetic-tts-en.multilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.multilingual-synthetic-tts-es
multilingual-synthetic-tts-es
Mirror of the exact bmcore v24 holdout subset. Source: Reubencf/multilingual-synthetic-tts.
Local count agrees with published language count. Source card states reference-speaker consent and research/noncommercial use only; preserve restriction.
Contains 8000 synthetic audio files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: 332a32ddb6ff09935516cf5bf24a7fb0a5ec066c.
Original source… See the full description on the dataset page: https://huggingface.co/datasets/34data/multilingual-synthetic-tts-es.Malian-multilingual-SNAC-TTS-dataset-vvmerged_multilingual_tts_6k_unsloth
