multilingual-tts
open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.Multilingual-TTS
Multilingual-TTS
A large multilingual corpus for pretraining TTS/STT models, gathered and normalized from 230+ public sources. ~191k hours of audio across 150+ languages, organized into 1,544 dataset configs and tokenized to 34.5B NeuCodec speech tokens (50 Hz, single-codebook) over 111.1M clips.
Each config is one source dataset normalized to rows of {audio_filename, text, speaker}:
audio_filename — clip path inside that config's <config>_audio.zip (mono MP3).
text —… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS.TTS-Multilingual-Test-Set
Overview
To assess the multilingual zero-shot voice cloning capabilities of TTS models, we have constructed a test set encompassing 24 languages. This dataset provides both audio samples for voice cloning and corresponding test texts.
Specifically, the test set for each language includes:
100 distinct test sentences.
Audio samples from two speakers (one male and one female) carefully selected from the Mozilla Common Voice (MCV) dataset, intended for voice cloning.
Researchers can… See the full description on the dataset page: https://huggingface.co/datasets/MiniMaxAI/TTS-Multilingual-Test-Set.Multilingual-TTS-language
Multilingual-TTS-language
malaysia-ai/Multilingual-TTS
with two extra columns:
column
description
audio_filename, text, speaker
unchanged from malaysia-ai/Multilingual-TTS
language
language detected from the text (transcript) column of every row
post-normalized
text after rule-based punctuation / capitalization normalization
All original columns and the file/folder layout are preserved: 1552 subsets /
1561 parquet files, 121,842,371 rows (August 2026). Every… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Multilingual-TTS-language.qwen3-tts-multilingual-emotional-speechne-tts-coqui-multilingual
NE-TTS Coqui Multilingual
Multilingual TTS dataset for 15 North East Indian languages, formatted for Coqui-AI VITS multilingual training. Contains 61,943 clips / 83.5 hours at 22050Hz (SNR >= 20dB only).
Languages
ISO
Language
Clips
Hours
grt
Garo
24,772
29.6
ccp
Chakma
10,689
14.3
nag
Nagamese
9,688
14.5
lus
Mizo
8,554
14.5
nnp
Wancho
5,081
6.3
trp
Kokborok
1,237
1.7
clk
Idu Mishmi
602
0.7
mjw
Karbi
373
0.4
nre
Rengma
258
0.4
nri… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-coqui-multilingual.
