datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yodas-granary
Dataset Card for YODAS-Granary
Repository: NeMo-speech-data-processor: Granary
Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages
Shared by: ESPnet
Dataset Description
YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.yodas2_sidon
YODAS2-Sidon
Overview
This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling.
YODAS-2 is a massive, multilingual YouTube-derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high-quality generation tasks.
We resampled original sidon output to 24kHz due to a storage constraints.
The dataset is provided in… See the full description on the dataset page: https://huggingface.co/datasets/sarulab-speech/yodas2_sidon.yodas_owsmv4🏆 News: Our OWSM v4 paper won the Best Student Paper Award at INTERSPEECH 2025!
Dataset Card for YODAS_OWSMv4
Paper: OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning (Best Student Paper at INTERSPEECH 2025)
Authors: Yifan Peng, Muhammad Shakeel, Yui Sudo, William Chen, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe
Data Cleaning Scripts: ESPnet
Model Demo: Gradio
Dataset Description
Open Whisper-style Speech Model (OWSM)is the first… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas_owsmv4.emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
yodas-ja000
YODAS Japanese (ja000)
Japanese manual caption subset of the YODAS dataset, repackaged for easier use.
Source
Original dataset: espnet/yodas (ja000 config)
Paper: YODAS: YouTube-Oriented Dataset for Audio and Speech
License: CC BY 3.0
Citation
If you use this dataset, please cite the original YODAS paper:
yodas2
YODAS2 for 🇺🇦 Ukrainian
Ukrainian validated subset of YODAS2
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 400213
Total duration: 998h 41m 3s
yodas2_sidon_th_tts
Thai TTS Dataset — Filtered & Quality-Verified from YODAS2 sidon
A filtered, quality-verified Thai text-to-speech dataset derived from sarulab-speech/yodas2_sidon, with transcriptions verified by multiple ASR models and Gemini, text fully normalized to Thai, and audio quality-screened with DNSMOS.
Dataset Summary
Samples
141,927
Audio hours
156.0
Speakers
4,199
Sample rate
24,000 Hz
Format
WAV, PCM 16-bit, mono
Language
Thai
Source… See the full description on the dataset page: https://huggingface.co/datasets/Chalermdej/yodas2_sidon_th_tts.yodas3
YODAS v3
Paper
YODAS v3 is a large web-crawled dataset containing over 1.1 million hours of audio that were originally released under a CC-BY-3.0 license. The dataset contains audio in over 100 languages. YODAS v3 can be used for a variety of multi-modal tasks, including Automatic Speech Recognition, Text-to-Speech, and Audio Representation Learning. We crawl a distinct set of videos from the v1 and v2 versions of YODAS, to guarantee that there are no overlaps in the data.
For… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas3.YodasSpeakerPoolUse this dataset in conjuction with:
https://github.com/fangningshao/YodasSpeakerPool
YodasSpeakerPool
YodasSpeakerPool is a curated, richly-annotated multi-speaker dataset featuring 7,600 unique speakers (3.4K Chinese, 4.2K English).
Derived from the Emilia-YODAS corpus, each sample is annotated by Gemini 2.5 Flash for its vocal characteristics and audio quality.
Dataset Features
Audio Specs: 4–15 second WAV samples of clean speech.
Rich Metadata: Includes ASR… See the full description on the dataset page: https://huggingface.co/datasets/HuHaiYang/YodasSpeakerPool.YodasSpeakerPoolUse this dataset in conjuction with:
https://github.com/fangningshao/YodasSpeakerPool
YodasSpeakerPool
YodasSpeakerPool is a curated, richly-annotated multi-speaker dataset featuring 7,600 unique speakers (3.4K Chinese, 4.2K English).
Derived from the Emilia-YODAS corpus, each sample is annotated by Gemini 2.5 Flash for its vocal characteristics and audio quality.
Dataset Features
Audio Specs: 4–15 second WAV samples of clean speech.
Rich Metadata: Includes ASR… See the full description on the dataset page: https://huggingface.co/datasets/fangningshao/YodasSpeakerPool.emilia-yodas-en-aligned
Emilia-YODAS EN Word-Aligned
Word-level forced-alignment timestamps for the English subset of
amphion/Emilia-Dataset
(Emilia-YODAS split), produced with
Qwen/Qwen3-ForcedAligner-0.6B.
No audio is redistributed — this dataset contains only metadata (IDs,
transcripts already present in Emilia-YODAS, and per-word [start, end]
timestamps). To use it, join on id with the original Emilia-YODAS audio.
Stats
Metric
Value
Utterances
4,516,833
Total audio
11,572.7… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-aligned.yodas2-opus
YODAS2 for 🇺🇦 Ukrainian (OPUS)
Ukrainian validated subset of YODAS2
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 400213
Total duration: 998h 41m 3s
yodas-en-replay
YODAS-EN Replay
General-domain English speech for replay mixing during domain adaptation, with
cased and punctuated transcripts. Audio comes from
espnet/yodas2 (CC-BY-3.0, sourced
from Creative-Commons YouTube videos); the transcripts are our own, produced with
faster-whisper large-v3-turbo.
Why this exists
If you fine-tune a small ASR model on a narrow domain, it forgets everything else.
We measured 5 hours of meeting audio buying 0.74 WER points in-domain while… See the full description on the dataset page: https://huggingface.co/datasets/moonshine-ai/yodas-en-replay.JA_Emilia_Yodas_ScribeEvents
JA Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/JA_Emilia_Yodas_266h containing only samples with ElevenLabs Scribe v1 audio events.
Changes from source
Filtered to rows where events_scribe is non-empty (4433 rows kept)
Bracket format unified: (event) in text_scribe replaced with [event]
Event types include
Vocal bursts: <laughs>, <sighs>, <clears throat>, etc.
Background: <background noise>, <music>, etc.
Other: <pause>, <unintelligible>… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/JA_Emilia_Yodas_ScribeEvents.distilled-yodas-spanishDistilled YODAS Spanish is a high-quality subset of the Spanish portion of the YouTube-Oriented Dataset for Audio and Speech (YODAS). While the full YODAS corpus contains over 37,000 hours of Spanish speech across 43 million files, this dataset provides a distilled version of approximately 8,000 validated hours.yodas_br
Description
Partie en breton du jeu de données espnet/yodas plus précisément les audios sont en bretons et le texte peut être en breton ou en français. D'après les auteurs, les textes et audios sont issus de vidéos YouTube sous licence CC.Les 388 premières lignes du jeu de données ont des textes en breton et représentent 14min et 57s. Les lignes 491 lignes suivantes ont des textes en français et représentent 23min et 1s.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Bretagne/yodas_br.EN_Emilia_Yodas_ScribeEvents
EN Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/EN_Emilia_Yodas_616h containing only samples with ElevenLabs Scribe v1 audio events (vocal bursts, background sounds, etc.).
Changes from source
Filtered to only include rows where events_scribe is non-empty (16017 rows out of 228,265 original)
Bracket format unified: Round brackets (laughs) in text_scribe replaced with square brackets [laughs] for consistency with vocal burst annotation format… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/EN_Emilia_Yodas_ScribeEvents.DE_Emilia_Yodas_ScribeEvents
DE Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/DE_Emilia_Yodas_680h containing only samples with ElevenLabs Scribe v1 audio events.
Changes from source
Filtered to rows where events_scribe is non-empty (12173 rows kept)
Bracket format unified: (event) in text_scribe replaced with [event]
Event types include
Vocal bursts: <laughs>, <sighs>, <clears throat>, etc.
Background: <background noise>, <music>, etc.
Other: <pause>, <unintelligible>… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/DE_Emilia_Yodas_ScribeEvents.espnet_yodas2Ce répertoire est vide, il a été créé pour améliorer le référencement du jeu de données espnet/yodas2.
espnet_yodasCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données espnet/yodas.
espnet_yodas_owsmv4Ce répertoire est vide, il a été créé pour améliorer le référencement du jeu de données espnet/yodas_owsmv4.
