datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yodas-granary
Dataset Card for YODAS-Granary
Repository: NeMo-speech-data-processor: Granary
Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages
Shared by: ESPnet
Dataset Description
YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.Bagpiper_SFT_Data
Bagpiper SFT Data
Release status: the validated Parquet release is being uploaded. The
homepage and metadata may appear before every large shard is committed.
Bagpiper SFT Data is the supervised fine-tuning corpus for
Bagpiper, an open-ended audio language model
that understands and generates speech, music, environmental sound, and their
mixtures through rich textual captions and planning.
The public release has exactly two configurations:
Configuration
Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.floras
FLORAS
FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language.
The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models.
Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers.
To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.yodas_owsmv4🏆 News: Our OWSM v4 paper won the Best Student Paper Award at INTERSPEECH 2025!
Dataset Card for YODAS_OWSMv4
Paper: OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning (Best Student Paper at INTERSPEECH 2025)
Authors: Yifan Peng, Muhammad Shakeel, Yui Sudo, William Chen, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe
Data Cleaning Scripts: ESPnet
Model Demo: Gradio
Dataset Description
Open Whisper-style Speech Model (OWSM)is the first… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas_owsmv4.Bagpiper_TTS_SFT_Data
Bagpiper-TTS SFT Data
Release status: the validated Parquet release is being uploaded. The
homepage and metadata may appear before every large shard is committed.
Bagpiper-TTS SFT Data supports
Bagpiper-TTS, a universal
speech-synthesis model that interprets free-form natural-language requests,
plans the requested delivery, produces a rich textual caption, and synthesizes
the target audio.
The release is organized into the six applications used by the paper:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_TTS_SFT_Data.Bagpiper_PreTrain_Data
Bagpiper Pretraining Data
Bagpiper Pretraining Data is the public rich-captioned audio snapshot associated
with Bagpiper, an open-ended audio language
model that learns bidirectional mappings between audio and comprehensive text
descriptions across speech, music, environmental sound, and mixtures.
The en metadata describes the primary rich-caption language. Source audio can
contain speech or singing in other languages; it is not an English-only audio
guarantee.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_PreTrain_Data.ace-opencpop-segments
Citation Information
@misc{shi2024singingvoicedatascalingup,
title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing},
author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe},
year={2024},
eprint={2401.17619},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2401.17619},
}
mms_ulab_v2MMS ulab v2 is a a massively multilingual speech dataset that contains 8900 hours of unlabeled speech across 4023 languages. In total, it contains 189 language families.
It can be used for language identification, spoken language modelling, or speech representation learning.
MMS ulab v2 is a reproduced and extended version of the MMS ulab dataset originally proposed in Scaling Speech Technology to 1000+ Languages, covering more languages and containing more data.
This dataset includes the raw… See the full description on the dataset page: https://huggingface.co/datasets/espnet/mms_ulab_v2.raw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16kLichessGamesraw_tts_esc_ESPnet_espnet_mls-multi_soundstream_16klol-esports-matches
GPTilt: League of Legends Esports Matches
This dataset is part of the GPTilt open-source initiative, aimed at democratizing access to high-quality LoL data for research and analysis, fostering public exploration, and advancing the community's understanding of League of Legends through data science and AI. It provides a clean, canonical record of the competitive matches and games of professional League of Legends.
By using this dataset, users accept full responsibility for any… See the full description on the dataset page: https://huggingface.co/datasets/gptilt/lol-esports-matches.code_search_net_python_10000_examplesESpeech-webinars2
Webinar Audio Dataset
Dataset Description
This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Asessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.ace-kising-segments
Citation Information
@misc{shi2024singingvoicedatascalingup,
title={Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing},
author={Jiatong Shi and Yueqian Lin and Xinyi Bai and Keyi Zhang and Yuning Wu and Yuxun Tang and Yifeng Yu and Qin Jin and Shinji Watanabe},
year={2024},
eprint={2401.17619},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2401.17619},
}
EsportsBench
EsportsBench: A Collection of Datasets for Benchmarking Rating Systems in Esports
EsportsBench is a collection of 20 esports competition datasets. Each row of each dataset represents a match played between either two players or two teams in a professional video game tournament.
The goal of the datasets is to provide a resource for comparison and development of rating systems used to predict the results of esports matches based on past results. Date is complete up to 2026-03-31.… See the full description on the dataset page: https://huggingface.co/datasets/EsportsBench/EsportsBench.raw_tts_esc_ESPnet_espnet_mls-english_soundstream_16kDSUChallenge2024
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
Paper: https://www.isca-archive.org/interspeech_2024/chang24b_interspeech.html
Arxiv: https://arxiv.org/abs/2406.07725
Challenge details: https://www.wavlab.org/activities/2024/Interspeech2024-Discrete-Speech-Unit-Challenge/
To cite:
@inproceedings{chang24b_interspeech,
title = {The Interspeech 2024 Challenge on Speech Processing Using Discrete Units},
author = {Xuankai Chang and Jiatong Shi and… See the full description on the dataset page: https://huggingface.co/datasets/espnet/DSUChallenge2024.lol-esports-entities
GPTilt: League of Legends Esports Directory
This dataset is part of the GPTilt open-source initiative, aimed at democratizing access to high-quality LoL data for research and analysis, fostering public exploration, and advancing the community's understanding of League of Legends through data science and AI. It provides a clean, canonical reference for the people and organizations of competitive League of Legends.
By using this dataset, users accept full responsibility for any… See the full description on the dataset page: https://huggingface.co/datasets/gptilt/lol-esports-entities.wikitonguesThe WikiTongues speech corpus is a collection of conversational audio across 700+ languages.
It can be used for spoken language modelling or speech representation learning.
This dataset includes the raw unsegmented audio in a 16kHz single channel format.
Each clip is usually 2-10 minutes long, and contains one or more speakers conversing in their language(s).
Sometimes, a speaker may switch languages within a single clip.
The total dataset size is around 70 hours.
The current version of the… See the full description on the dataset page: https://huggingface.co/datasets/espnet/wikitongues.ml_superb_hfgo2-air-controlbench-v1
Go2 Air ControlBench v1 Public Preview
Go2 Air ControlBench is a compact command-to-outcome benchmark for a stock Unitree Go2 Air. It asks a narrow question that matters for robot planners and world-model scorers:
If we command the robot to move, what actually happens, and which candidate command should a planner have selected?
The release is deliberately small and claim-bounded. It is not an imitation-learning corpus and not mocap-grade ground truth. It is a public-safe… See the full description on the dataset page: https://huggingface.co/datasets/espejelomar/go2-air-controlbench-v1.espeech_balalaika
ESpeech datasets (w/o podcasts) Annotated by Balalaika
[!IMPORTANT]
Official dataset for our INTERSPEECH 2026 paper
"A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models" (arXiv:2507.13563).
Part of the Balalaika Russian speech data-processing pipeline — code: https://github.com/lab260ru/balalaika.
If you use this resource, please cite it.
A curated Russian speech dataset for advanced speech generative tasks.… See the full description on the dataset page: https://huggingface.co/datasets/lab260/espeech_balalaika.VidTouch
VidTouch
VidTouch is a specimen-level visuo-tactile benchmark for material understanding
and material knowledge discovery. It links repeated RGB observations and DIGIT
tactile videos of the same physical fabric to four semantic axes:
weave;
ordered material composition;
usage;
functional features.
We purchased 144 physical fabric swatches and independently acquired all released RGB images and tactile videos.
We used the supplier descriptions as source records. As they were… See the full description on the dataset page: https://huggingface.co/datasets/AQUILA-espresso/VidTouch.jesus_dramasJesus Dramas is a collection of religious audio dramas across 430 languages. In total, there is around 640 hours of audio.
It can be used for language identification, spoken language modelling, or speech representation learning.
This dataset includes the raw unsegmented audio in a 16kHz single channel format. Each audio drama can have multiple speakers, for both male and female voices.
It can be segmented into utterances with a voice activity detection (VAD) model such as this one.
The… See the full description on the dataset page: https://huggingface.co/datasets/espnet/jesus_dramas.long-yodas-unsegmentedJobOffers_ESP
Ofertas de Empleo Públicas en España (EURES, 2025)
https://doi.org/10.57967/hf/6740
Dataset Summary
Este dataset recopila ofertas de empleo publicadas en portales oficiales de empleo europeos y españoles, principalmente a través de la red EURES (European Employment Services).Forma parte del proyecto desarrollado para una práctica en la asignatura Descubrimiento del Conocimiento en Datos Complejos del grado de Ciencia de Datos e Inteligencia Artificial en la Universidad… See the full description on the dataset page: https://huggingface.co/datasets/MiguelGP-13/JobOffers_ESP.ESpeech-tuchniyzhab
Tuchniy Zhab YouTube Audio Dataset
Dataset Description
This dataset contains 306 hours of processed audio segments extracted from the "Tuchniy Zhab" YouTube channel with corresponding metadata. Each audio file represents a segment from the channel's videos and content, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Task: TTS, ASR, Quality Assessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON metadata… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-tuchniyzhab.esperanto-mt-parallel-v13
esperanto-mt-parallel v13
EN<->EO parallel training corpus. 5,025,333 rows after dedup.
What changed vs v12 (jensjepsen/esperanto-mt-parallel)
Dropped Helsinki-NLP/opus-100 en-eo train split (144,549 rows).
It is an aggregated multilingual blob that bundles KDE4/GNOME/Ubuntu
.po localization pairs without src labels. In v12 this caused a
systematic MT failure mode: capitalized-fragment-no-terminal-punct
inputs collapsed to memorized UI labels
(e.g. "@ info:… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/esperanto-mt-parallel-v13.ESpeech-podcasts
Podcasts Audio Dataset
Dataset Description
This dataset contains 3200 hours of processed audio segments extracted from various podcasts with corresponding metadata. Each audio file represents a segment from podcast episodes, processed at 44.1kHz sample rate.
Dataset Summary
Language: Russian
Total Duration: 3200 hours of speech
Task: TTS, ASR, Quality Assessment
Audio format: MP3, 44.1kHz sample rate
Structure: Segmented audio files with JSON… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-podcasts.
