datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
libritts_r_filtered
Dataset Card for Filtered LibriTTS-R
This is a filtered version of LibriTTS-R. It has been filtered based on two sources:
LibriTTS-R paper [1], which lists samples for which speech restoration have failed
LibriTTS-P [2] list of excluded speakers for which multiple speakers have been detected.
LibriTTS-R [1] is a sound quality improved version of the LibriTTS corpus which is a multi-speaker English corpus of approximately
585 hours of read English speech at 24kHz sampling rate… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/libritts_r_filtered.ramanv-tts-all-raw
ramanv-tts-all-raw
Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages.
emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
open-bible
OpenBibleTTS
OpenBibleTTS is a large-scale, multilingual speech corpus for low-resource text-to-speech (TTS), spanning 37 underrepresented languages across five regions. It contains ~3,469 hours of aligned, verse-level read speech and 1,121,956 utterances, derived from the Open Bible platform and released under a permissive license.
Alignment pipeline: https://github.com/davidguzmanr/open-bible-resources
Source: Open Bible (CC BY-SA)
Languages
Africa (19), South… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-tts/open-bible.mls_eng
Dataset Card for English MLS
Dataset Summary
This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng.tts_farm
Multilingual TTS/ASR Aggregated Dataset
Cleaned, deduplicated and loudness-normalized Arabic, Japanese, Korean, Turkish, and Vietnamese speech. The training columns are audio (16-bit PCM WAV, 22050 Hz) and text; the remaining columns contain quality and provenance metadata.
majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.mls_eng_10k
Dataset Summary
This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls_eng_10k.commonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.parler-tts_mls_eng_10k_snac_token_old
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/parler-tts_mls_eng_10k_snac_token_old.ewe-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
48775 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-bible-audio-text-tts.WolneLektury-TTS-Polish
WolneLektury-TTS-Polish
A large-scale, high-quality Polish speech dataset for text-to-speech and automatic speech recognition.
Data Source
Derived from Wolne Lektury (Free Readings), a Polish digital library with public domain audiobooks featuring professional voice actors.
Dataset Statistics
Metric
Value
Total samples
383,710
Total duration
997 hours
Unique narrators
1207
Male samples
294,756 (767h)
Female samples
88,945 (230h)
Average… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/WolneLektury-TTS-Polish.TTSDistil-Phonology
Overview
This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training.
The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization.
This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil-Phonology.cantonese-youtube-tts
Cantonese Audio TTS Dataset
This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered towards TTS (text-to-speech) use cases, more than the 2 previously published datasets, as there is more extensive filtering and audio enhancement. For speaker labelling, you can use speaker embedding models like Nvidia's TitaNet
Filtered out:
Overlapped voices, detected using pyannote/speaker-diarization-3.1
Music, detected using a… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/cantonese-youtube-tts.VieNeu-TTS-140h
pnnbao-ump/VieNeu-TTS-140h
Mô tả Dataset
A high-quality Vietnamese Text-to-Speech (TTS) dataset containing 74,858 audio samples with phonemized transcripts. This benchmark dataset is designed for fine-tuning modern TTS models with maximum synthesis quality. The text corpus is completely phonemized using standard international phonetic alphabet (IPA) representations suitable for neural acoustic modeling.
Quick Facts
Language: Vietnamese 🇻🇳
Tasks:… See the full description on the dataset page: https://huggingface.co/datasets/pnnbao-ump/VieNeu-TTS-140h.ghana-named-entities-tts-twi
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghana Named Entities TTS — Twi
A Twi-language speech dataset built from descriptions of Ghana named entities
(people, places, organisations, and concepts). Each audio clip is a synthesised
reading of a passage that describes several… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-named-entities-tts-twi.dagbani-bible-audio-text-tts
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
53410 speech-text pairs split from long recordings.
Processing pipeline
Source audio from ghananlpcommunity/dagbani-tts-bible-full-audio-text
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/dagbani-bible-audio-text-tts.TTSDistil2
Overview
This repository contains a speech dataset developed for Text-to-Speech (TTS) research and model training.
The dataset is part of an ongoing research project focused on building high-quality speech corpora for modern neural TTS systems. It is actively maintained, with continuous improvements in data quality, transcription consistency, metadata, and organization.
This repository is intended to host research data used throughout the development process and is not intended… See the full description on the dataset page: https://huggingface.co/datasets/ShiniChien/TTSDistil2.bible_tts_hausa
Dataset Card for BibleTTS Hausa
Dataset Summary
BibleTTS is a large high-quality open Text-to-Speech dataset with up to 80 hours of single speaker, studio quality 48kHz recordings.
This is a Hausa part of the dataset. Aligned hours: 86.6, aligned verses: 40,603.
Languages
Hausa
Dataset Structure
Data Fields
audio: audio path
sentence: transcription of the audio
locale: always set to ha
book: 3-char book encoding
verse: verse id… See the full description on the dataset page: https://huggingface.co/datasets/vpetukhov/bible_tts_hausa.TTS-German
TTS-German
High-quality German speech dataset for TTS and ASR, derived from CML-TTS German.
Processing Pipeline
Standardize → 24kHz mono WAV, loudness normalize
Transcribe → WhisperX word-level timestamps
Segment → ≤12s at word boundaries
Denoise → DeepFilterNet
Quality filter → DNSMOS ≥ 2.5
G2P → IPA phonemes (custom dictionary)
Statistics
Metric
Value
Samples
670,509
Hours
1250h
Sample rate
24kHz mono
Max duration
12s
Schema… See the full description on the dataset page: https://huggingface.co/datasets/datadriven-company/TTS-German.tts-rj-hi-karya
Rajasthani Hindi Speech Dataset
This dataset consists of audio recordings of participants reading out stories in Rajasthani Hindi, one sentence at a time. They had 98 participants from Soda, Rajasthan. Each participant read 30 stories. In total, we have 426872 recordings in this dataset. They had roughly 58 male participants and 40 female participants.
Point to Note:
While random sampling suggests that most users have to their best effort tried to accurately read out the sentences… See the full description on the dataset page: https://huggingface.co/datasets/1rsh/tts-rj-hi-karya.twi-tts-asr
Ghana Twi Speech Dataset (TTS + ASR)
79,655 synthetic Twi speech samples (88.8 hours) generated with OmniVoice. Designed for both text-to-speech (TTS) and automatic speech recognition (ASR) research on Twi.
Intended use
TTS / Voice cloning: Use the audio + text pairs directly. The dataset includes diverse voice profiles for building or fine-tuning Twi TTS systems.
ASR: Use the text transcripts as labels. The audio is natively at 24 kHz — resample to 16 kHz for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-tts-asr.Turkish_TTS_Datamls-eng-speaker-descriptions
Dataset Card for Annotations of English MLS
This dataset consists in annotations of the English subset of the Multilingual LibriSpeech (MLS) dataset.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K hours for other languages.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/parler-tts/mls-eng-speaker-descriptions.yodas2_sidon_th_tts
Thai TTS Dataset — Filtered & Quality-Verified from YODAS2 sidon
A filtered, quality-verified Thai text-to-speech dataset derived from sarulab-speech/yodas2_sidon, with transcriptions verified by multiple ASR models and Gemini, text fully normalized to Thai, and audio quality-screened with DNSMOS.
Dataset Summary
Samples
141,927
Audio hours
156.0
Speakers
4,199
Sample rate
24,000 Hz
Format
WAV, PCM 16-bit, mono
Language
Thai
Source… See the full description on the dataset page: https://huggingface.co/datasets/Chalermdej/yodas2_sidon_th_tts.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.turkish-tts-audiobooks
Turkish TTS Audiobooks
Turkish read-speech corpus for text-to-speech training, built from Turkish
audiobook and spoken-article recordings by an automatic pipeline: VAD
segmentation → technical QC → acoustic event tagging → DNSMOS → speaker
embedding/consistency → double-pass Whisper ASR → text policy → leakage-free
splitting. Audio is 16 kHz mono lossless FLAC embedded in the Parquet shards.
The pipeline that produced it — every stage, every threshold, the export and
audit… See the full description on the dataset page: https://huggingface.co/datasets/serdarcaglar/turkish-tts-audiobooks.indic_tts_ml
Indic TTS Malayalam Speech Corpus
The Malayalam subset of Indic TTS Corpus, taken from
this Kaggle database. The corpus contains
one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given
in the repository.
vieneu-tts-140h-dataset
pnnbao-ump/VieNeu-TTS-140h
Mô tả Dataset
Dataset tiếng Việt chất lượng cao cho Text-to-Speech (TTS) với 74,858 mẫu audio và transcript được phonemize.
Mục tiêu của mình là tạo bộ dataset chuẩn mực để finetune các model TTS hiện nay với chất lượng cao nhất. Mình thu thập audio chất lượng cao từ youtube, làm sạch nền, loại bỏ noise, dùng whisper-large-v3 để tạo transcription, sau đó cho Agent sửa lỗi chính tả và feedback lại cho con người. Bộ dữ liệu cũng được phonemize hóa… See the full description on the dataset page: https://huggingface.co/datasets/LanguaMan/vieneu-tts-140h-dataset.
