datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EgoIT-99KCheckout the paper EgoLife (https://arxiv.org/abs/2503.03803) for more information.
WSC-EvalFMSU-BenchExisting benchmarks predominantly cater to macroscopic tasks and suffer from coarse annotation granularity. We construct FMSU-Bench, a pioneering Fine-grained Multi-dimensional Speech Understanding Benchmark.
Scale & Scope: Comprises over 24,000 bilingual instances (Chinese/English), manually verified by domain experts.
Comprehensive Taxonomy: Systematically covers 14 distinct speech dimensions structured into a 5-tier taxonomy:
Speaker Demographics: Gender, Age, Accent
Acoustic-Prosodic… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/FMSU-Bench.HumDial-EIBench
HumDial-EIBench
A human-recorded multi-turn emotional intelligence benchmark for audio language models.
Figure: Three-stage data pipeline and the four evaluation tasks in HumDial-EIBench.
HumDial-EIBench is designed to evaluate whether audio language models (ALMs) truly understand emotion in speech, rather than relying on text transcription shortcuts.
The benchmark is built from authentic human-recorded dialogues from the ICASSP 2026 HumDial Challenge and… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/HumDial-EIBench.voicebench
License
The dataset is available under the Apache 2.0 license.
Citation
If you use the VoiceBench dataset in your research, please cite the following paper:
@article{chen2024voicebench,
title={VoiceBench: Benchmarking LLM-Based Voice Assistants},
author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou},
journal={arXiv preprint arXiv:2410.17196},
year={2024}
}
VoiceAgentBench
VoiceAgentBench
This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978).
VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.QuranTTS
QuranTTS v4
An ear-verified Quranic recitation corpus for TTS and speech restoration.
Built in-house for the lab's own training runs. We have since moved on to a larger corpus and a newer
pipeline, so this one is published rather than shelved. What you get is the corpus exactly as it stood when
we stopped using it: complete, documented, and unmaintained.
Non-commercial, strictly. Neither this corpus nor any model trained on it may be used for any
commercial purpose, and… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/QuranTTS.fma-labeled
FMA Labeled — Multi-Attribute Music Dataset
🏆 Submitted to the Uncharted Data Challenge
hosted by Adaption Labs — credit to
Adaptive Data by Adaption for organizing the hackathon.
A large-scale labeled music dataset built on top of the Creative-Commons
subset of the Free Music Archive (FMA). Every
track has been automatically annotated with lyrics, genre, mood, instruments,
tempo, key, and more using Google Gemini (gemini-flash-latest).
Intended for training and evaluating music… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/fma-labeled.JamendoMaxCaps
JamendoMaxCaps Dataset
JamendoMaxCaps is a large-scale dataset of over 362,000 instrumental tracks sourced from the Jamendo platform. It includes generated music captions and original metadata. Additionally, we introduce a retrieval system that utilizes both musical features and metadata to identify similar songs, which are then used to impute missing metadata via a local large language model (LLLM).
This dataset facilitates research in:
Music-language understanding
Music… See the full description on the dataset page: https://huggingface.co/datasets/amaai-lab/JamendoMaxCaps.MMAU-Pro
MMAU-Pro: A Challenging and Comprehensive Benchmark for Audio General Intelligence
MMAU-Pro is the most comprehensive benchmark to date for evaluating audio intelligence in multimodal models. It spans speech, environmental sounds, music, and their combinations—covering 49 distinct perceptual and reasoning skills.
The dataset contains 5,305 expert-annotated question–answer pairs, with audios sourced directly from the wild. It introduces several novel challenges overlooked by… See the full description on the dataset page: https://huggingface.co/datasets/gamma-lab-umd/MMAU-Pro.SongEval
SongEval 🎵
A Large-Scale Benchmark Dataset for Aesthetic Evaluation of Complete Songs
📖 Overview
SongEval is the first open-source, large-scale benchmark dataset designed for aesthetic evaluation of complete songs. It provides over 2,399 songs (~140 hours) annotated by 16 expert raters across five perceptual dimensions. The dataset enables research in evaluating and improving music generation systems from a human aesthetic perspective.
🌟 Features… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/SongEval.mmauUrduSpeech
Dataset Summary
UrduSpeech is a large-scale, high-fidelity Urdu speech corpus comprising 156 hours of audio with comprehensive 12-dimensional paralinguistic metadata. The corpus addresses the critical under-resourcing of Urdu in speech technology by providing:
71,792 diarized utterances across diverse content categories
Three specialized subsets: Standard Pakistani Urdu (US-Std, 59.2h), Urdu-English Code-Switched (US-CS, 89.4h), and Pakistani-Accented English (US-EngPk, 7.3h)… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/UrduSpeech.wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.audioset_cla_label_des_naivelibrispeechMMAU-test-miniSonicMasterDatasetThe SonicMaster dataset is a large collection of paired degraded and high-quality music tracks, introduced in the paper SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering.
This dataset was constructed by applying nineteen degradation functions belonging to five enhancement groups: equalization, dynamics, reverb, amplitude, and stereo. It is designed to train unified generative models for music restoration and mastering. The original music files were sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/amaai-lab/SonicMasterDataset.voicebench
License
The dataset is available under the Apache 2.0 license.
Citation
If you use the VoiceBench dataset in your research, please cite the following paper:
@article{chen2024voicebench,
title={VoiceBench: Benchmarking LLM-Based Voice Assistants},
author={Chen, Yiming and Yue, Xianghu and Zhang, Chen and Gao, Xiaoxue and Tan, Robby T. and Li, Haizhou},
journal={arXiv preprint arXiv:2410.17196},
year={2024}
}
OSSL-v2
Open Screen Soundtrack Libary Version 2 (OSSL-v2)
Paired film video ↔ soundtrack music clips for video-to-music generation.
Paper: Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Layout
ossl-v2-hf/
├── metadata.csv # one row per movie (film + source metadata)
├── splits/{train,test}.txt # clip_ids per split
├── public_train_test_remapped.pkl # {"train":[clip_id...], "test":[clip_id...]}
├── train/{video… See the full description on the dataset page: https://huggingface.co/datasets/McAuley-Lab/OSSL-v2.WenetSpeech_tempBERSt
BERSt Dataset
We release the BERSt Dataset for various speech recognition tasks including Automatic Speech Recognition (ASR) and Speech Emotion Recogniton (SER)
Read the paper here
Overview
4526 single phrase recordings (~3.75h)
98 professional actors
19 phone positions
7 emotion classes
3 vocal intensity levels
varied regional and non-native English accents
nonsense phrases covering all English Phonemes
Data collection
The BERSt dataset represents data… See the full description on the dataset page: https://huggingface.co/datasets/Rosie-Lab/BERSt.WorldSensegigaspeechsam-wake-word-raw-datamelodySim
MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection
Github | Model | Paper
The MelodySim dataset contains 1,710 valid synthesized pieces originated from Slakh2100 dataset, each containing 4 different versions (through various augmentation settings), with a total duration of 419 hours.
This dataset may help research in:
Music similarity learning
Music plagiarism detection
Dataset Details
The MelodySim dataset contains three splits: train… See the full description on the dataset page: https://huggingface.co/datasets/amaai-lab/melodySim.labeled_datadutch-tts-labeled-complete
Dutch TTS Dataset - Complete Labeled
A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of audio data.
Quick Preview
The default config shows a 100-row sample for the dataset viewer. To access the full dataset, use the full config.
Dataset Description
This dataset contains Dutch speech recordings with rich metadata including:
Emotion labels (neutral, happy, sad, angry)
Speaker IDs (239,388 unique speakers)… See the full description on the dataset page: https://huggingface.co/datasets/AITRADER/dutch-tts-labeled-complete.WenetSpeech-Wu-Bench
WenetSpeech-Wu Bench
We introduce WenetSpeech-Wu-Bench, the first publicly available, manually curated benchmark for Wu dialect speech processing, covering ASR, Wu-to-Mandarin AST, speaker attributes, emotion recognition, TTS, and instruct TTS, and providing a unified platform for fair evaluation.
ASR: Wu dialect ASR (9.75 hour, including Shanghainese, Suzhounese, and Mandarin code-mixed speech). Evaluated by CER.
Wu→Mandarin AST: Speech translation from Wu dialects to Mandarin (3k… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/WenetSpeech-Wu-Bench.common_voice_16_1_hi_pseudo_labelled
