datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cml-tts
Dataset Card for CML-TTS
Dataset Summary
CML-TTS is a recursive acronym for CML-Multi-Lingual-TTS, a Text-to-Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG).
CML-TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/cml-tts.audiofolder_two_configs_in_metadatalibrispeech_asr_dummyaudiofolder_single_config_in_metadataMDCtasksThis dataset is for storing assets for https://huggingface.co/tasks and https://github.com/huggingface/huggingface.js/tree/main/packages/tasks
audiofolder_no_configs_in_metadatawhisper_transcriptions.reazon_speech_alltiny-testvoicehub-arena-seed-tts-eval
VoiceHub Arena — native TTS evaluations
Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA
speaker SIM and UTMOS22 measurements. The full campaign is still running.
Each generation method is evaluated separately using its publisher's native API.
Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target
diagnostic pilots are stored separately and must not be treated as full scores.
Interactive demo ·
Source code
Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.librivox-mirror
LibriVox Mirror
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Metric
Value
Published books
21,734
Published sections
493,396
Audio hours
132,613.0
Audio languages
86
Quarantined books
603
Last updated (UTC)
2026-09-24T13:43:06.550876Z
Audio by language
Language
Hours
English
131,663.6
German
417.0
Spanish
160.9
French
103.8
Portuguese
37.4
Polish
34.1
Dutch
25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.audiofolder_two_configs_in_metadatawhisper_transcriptions.mls.wer_10.0t2a-mommy
t2a-mommy
Female-voice ASMR corpus for the text2asmr project.
Previously published as aoxo/audios2.
Companion repos: aoxo/t2a-daddy (male voice),
aoxo/t2a-audios-v1 (the original v1 corpus).
Layout
path
what
<creator>/<title>.m4a
source audio, 48 kHz AAC, one folder per creator
<creator>/<title>.json
word-level Whisper large-v3 alignment ([] = skipped: near-silent or undecodable)
labels/qwen3omni.jsonl
non-speech ontology labels for gap clips… See the full description on the dataset page: https://huggingface.co/datasets/aoxo/t2a-mommy.dummy-audio-samplesaudiofolder_two_configs_in_metadata_with_defaultNatureLM-audio-training
Dataset card for NatureLM-audio-training
Overview
NatureLM-audio-training is a large and diverse audio-language dataset designed for training bioacoustic models that can generate a natural language answer to a natural language query on a reference bioacoustic audio recording.
For example, for an in-the-wild audio recording of a bird species, a relevant query might be "What is the common name for the focal species in the audio?" to which an audio-language model trained… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/NatureLM-audio-training.laions_got_talent
LAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel.
Dataset Composition
The dataset includes:
Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.esb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.TAGARELA
TAGARELA: A Portuguese Speech Dataset From Podcasts
TAGARELA is a large-scale Portuguese speech dataset built from podcast audio and curated for speech technology research, especially Automatic Speech Recognition (ASR) and Text-to-Speech (TTS).
The dataset contains more than 8,972 hours of Portuguese speech derived from the Cem Mil Podcasts collection. It includes Brazilian Portuguese and European Portuguese speech, processed through a pipeline involving audio standardization… See the full description on the dataset page: https://huggingface.co/datasets/freds0/TAGARELA.train-bn
Dataset Card for "train-bn"
More Information needed
tadabur-align-references
tadabur-align-references
Precomputed reference embeddings powering tadabur-align — word-level timestamp extraction for Quranic recitation via DTW alignment transfer (no ASR).
What this is
For 5,481 of the Quran's 6,236 ayahs, this dataset holds frame-level tadabur-embedding features for up to 8 reference reciters, plus each reference's word-level timestamps and internal-pause intervals. No audio is included — only model outputs and timing data. tadabur-align… See the full description on the dataset page: https://huggingface.co/datasets/FaisaI/tadabur-align-references.otoSpeech-full-duplex-turn-104h
Dataset Card for otoSpeech-full-duplex-turn-104h
Contact
Website: https://oto.earthEmail: agent@oto.earth
Dataset Summary
otoSpeech-full-duplex-turn-104h is an English, full-duplex conversational speech dataset for research on turn-taking and related spoken-dialogue phenomena. It contains 420 two-speaker conversations totaling approximately 104.94 hours. Each conversation includes time-aligned, channel-separated audio, a stereo combined recording… See the full description on the dataset page: https://huggingface.co/datasets/otoearth/otoSpeech-full-duplex-turn-104h.whisper_transcriptions.reazonspeech.alltelegram-audiobook-chizzled
Telegram Persian Audiobook Chizzled
1,555,434 Persian audiobook clips · 14,430.384 hours · 16 kHz mono PCM WAV · public Parquet release
This is a large, provenance-preserving collection of Persian audiobook audio gathered from 26 Telegram channels accessible to the collector account. Each source message is retained as message-level provenance and segmented with Silero voice-activity detection (VAD) into pause-aware clips. The audio bytes are embedded in Parquet files, so the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telegram-audiobook-chizzled.MultiMed-TSStest_librispeech_parquetmultilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.tat_youtubetraining-movies-test-stuff
