datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.qwen3-tts-multilingual-emotional-speechne-tts-coqui-multilingual
NE-TTS Coqui Multilingual
Multilingual TTS dataset for 15 North East Indian languages, formatted for Coqui-AI VITS multilingual training. Contains 61,943 clips / 83.5 hours at 22050Hz (SNR >= 20dB only).
Languages
ISO
Language
Clips
Hours
grt
Garo
24,772
29.6
ccp
Chakma
10,689
14.3
nag
Nagamese
9,688
14.5
lus
Mizo
8,554
14.5
nnp
Wancho
5,081
6.3
trp
Kokborok
1,237
1.7
clk
Idu Mishmi
602
0.7
mjw
Karbi
373
0.4
nre
Rengma
258
0.4
nri… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-coqui-multilingual.melo-tts-ghana-multilingual
Melo-TTS multilingual Ghana training dataset
Ghanaian multilingual TTS training data for a MeloTTS fine-tune, phonemised with
ghananlpcommunity/xlsr-twi-codeswitch-ipa.
Languages
twi-asantee — Asante Twi (AfriSpeech open-bible-speech-african)
twi-akuapem — Akuapem Twi (AfriSpeech open-bible-speech-african)
ewe — Ewe (AfriSpeech open-bible-speech-african)
twi-en-codeswitch — Twi-English code-switched speech (ghananlpcommunity)
ghana-english —… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/melo-tts-ghana-multilingual.multilingual-tts-voice-dataset
Multilingual TTS Voice Dataset
Multilingual speech and structured voice-control data for text-to-speech research and training.
The collection covers English, Kazakh, Kyrgyz, Russian, Tajik, Turkish, Turkmen, Uzbek, and mixed-language speech. Access requests are reviewed manually.
Configurations
audio: utterances with embedded audio.
prompt_specs: structured text and delivery specifications.
clone_pairs: same-speaker reference and target pairs.
voice_profiles:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/multilingual-tts-voice-dataset.multilingual-tts-benchmark
Multilingual Speech Benchmark for Zero-Shot TTS
A voice-cloning and intelligibility benchmark for six language variants, built
from Common Voice 17.0 by coverage-driven selection rather than random sampling.
Every example pairs a reference clip of one speaker with a target text that
speaker never read, so a system is asked to clone a voice and produce new
speech, which is what zero-shot TTS is actually for.
Pipeline source code:… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/multilingual-tts-benchmark.multilingual-tts
Before Anything and Everything ⚱
In the time of writing this Dataset Card, 17,490 18,412 civilian has been killed in Palestine (7,870 8,000 are children and 6,121 6,200 are women).
Seek any non-profit organization to help them with what you can (For myself, I use Mersal) 🇵🇸
Dataset Description
The Multilingual TTS dataset is an exceptional compilation of text-to-speech (TTS) samples, meticulously crafted to showcase the richness and diversity of human languages.… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/multilingual-tts.multilingual-synthetic-tts
Multilingual Synthetic TTS Dataset
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale synthetic multilingual speech dataset — 68,677 clips across
9 languages, generated with Qwen3-TTS-12Hz-1.7B-Base
using zero-shot voice cloning from 5 reference speakers.
Intended for training and evaluating TTS, ASR, voice conversion, and
multilingual speech models. Each clip is paired with the… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-synthetic-tts.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/YoonSeon/TTS-Multilingual-Test-Set.mossv1.5-tts-multilingual-test
mossv1.5-tts-multilingual-test
5 reference voices (spk_1..spk_5) -> EN/IT/JA/TH/ZH zero-shot clone, MOSS-TTS-v1.5. One row per voice, per-language audio/text/metric columns.
Lang
DNSMOS
spk_sim
n
en
3.82
0.93
5
it
3.93
0.80
5
ja
3.80
0.85
5
th
3.86
0.81
5
zh
3.80
0.83
5
Reference audio: Expresso (CC BY-NC 4.0, non-commercial).
multilingual-tts
Before Anything and Everything ⚱
In the time of writing this Dataset Card, 17,490 18,412 civilian has been killed in Palestine (7,870 8,000 are children and 6,121 6,200 are women).
Seek any non-profit organization to help them with what you can (For myself, I use Mersal) 🇵🇸
Dataset Description
The Multilingual TTS dataset is an exceptional compilation of text-to-speech (TTS) samples, meticulously crafted to showcase the richness and diversity of human languages.… See the full description on the dataset page: https://huggingface.co/datasets/huayidu139/multilingual-tts.multilingual-synthetic-tts-de
multilingual-synthetic-tts-de
Mirror of the exact bmcore v24 holdout subset. Source: Reubencf/multilingual-synthetic-tts.
Local count agrees with published language count. Source card states reference-speaker consent and research/noncommercial use only; preserve restriction.
Contains 8998 synthetic audio files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: 332a32ddb6ff09935516cf5bf24a7fb0a5ec066c.
Original source… See the full description on the dataset page: https://huggingface.co/datasets/34data/multilingual-synthetic-tts-de.TTS_Multilingual_Data
Documentation Dataset: TTS_Multilingual_Data
Dataset Summary
This large-scale multilingual corpus is designed for linguistic analysis and the development of speech processing models. It supports tasks such as Text-to-Speech (TTS), Automatic Speech Recognition (ASR), and speaker identification. Structured in Parquet format, it serves as a key resource for training and evaluating models, using metrics tailored to ASR and speech technologies.
Thematic Categories… See the full description on the dataset page: https://huggingface.co/datasets/Databoost/TTS_Multilingual_Data.multilingual-synthetic-tts-es
multilingual-synthetic-tts-es
Mirror of the exact bmcore v24 holdout subset. Source: Reubencf/multilingual-synthetic-tts.
Local count agrees with published language count. Source card states reference-speaker consent and research/noncommercial use only; preserve restriction.
Contains 8000 synthetic audio files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: 332a32ddb6ff09935516cf5bf24a7fb0a5ec066c.
Original source… See the full description on the dataset page: https://huggingface.co/datasets/34data/multilingual-synthetic-tts-es.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/jeshica/TTS-Multilingual-Test-Set.multilingual-tts-corpus
Multilingual TTS Corpus
A multilingual text-to-speech dataset containing audio recordings with text transcriptions across multiple languages. This is the first batch; more languages will be added over time.
Languages (Batch 1)
Subset
Language
Audio Files
Annotation Format
Source
ru-tts/
Russian
10
JSONL (single file)
俄语TTS(文本标注及语音)基石数据集 #294
ru-speech/
Russian
5
JSONL (single file)
俄语高质量语音音频语料库 #481
vi-speech/
Vietnamese
5
JSONL (single file)
越南语高质量语音音频语料库… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/multilingual-tts-corpus.multilingual-test-distill-strong-tts-20260520
Multilingual Test Distill Strong TTS 20260520
This repository contains a distributable tar-sharded version of multilingual_test_distill_strong_tts_20260520.
The dataset follows the local voice_dataset/data layout after extraction:
data/csvs/metadata_zh.csv
data/csvs/metadata_en.csv
data/csvs/metadata_ja.csv
data/csvs/metadata_ko.csv
data/zh/**/*.wav
data/en/**/*.wav
data/ja/**/*.wav
data/ko/**/*.wav
Metadata format:
file_path|duration|dnsmos|text
dnsmos is intentionally blank… See the full description on the dataset page: https://huggingface.co/datasets/guangzhaoli/multilingual-test-distill-strong-tts-20260520.merged_multilingual_tts_6k_unslothluel-multilingual-tts-samples
Multilingual TTS Samples (Luel)
License: All Rights Reserved. Proprietary. Access only for authorized parties; no redistribution or use without permission. See LICENSE.
A multilingual text-to-speech / read-speech dataset of short scripted utterances across 7 languages. Each sample is a single-speaker recording of a written prompt, paired with rich speaker and recording metadata. Useful for TTS training and evaluation, ASR adaptation, dialect/accent studies, and read-speech… See the full description on the dataset page: https://huggingface.co/datasets/Luel-ai/luel-multilingual-tts-samples.DataCatalyst_Multilingual_TTS_Sample
DataCatalyst Multilingual TTS Sample
Hindi | English | Hinglish (Code-Switched)Version: v1.2 | April 2026
Overview
This dataset is a multilingual speech sample curated by DataCatalyst for ASR and TTS evaluation.
Languages included:
Hindi
English (Indian accent)
Hinglish (code-switched)
Each language contains 6 utterances (~10 seconds each).
Key Features
48 kHz, 24-bit mono WAV
Speaker-consistent recordings
Loudness-normalized variants (EBU R128… See the full description on the dataset page: https://huggingface.co/datasets/DataCatalystAI/DataCatalyst_Multilingual_TTS_Sample.multilingual-music-tts-malayalam
