datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AudioJailbreak
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly.
📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.GSContextASR-Bench
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
Automatic Speech Recognition (ASR) has been extensively investigated, yet prior benchmarks have largely focused on assessing the acoustic robustness of ASR models, leaving evaluations of their linguistic capabilities relatively underexplored. This largely stems from the limited parameter sizes and training corpora of conventional ASR models, leaving them with insufficient world knowledge, which is crucial for… See the full description on the dataset page: https://huggingface.co/datasets/MrSupW/ContextASR-Bench.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/smgjch/meow-10k.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.speech-translation-and-summarization
English-Centric Multilingual Audio Dataset
This dataset contains generated article and summary audio for English-centric multilingual directions.
Each direction folder contains metadata JSONL files and corresponding audio files for few_shot and test splits.
Included directions
amharic_english / english_amharic
arabic_english / english_arabic
bengali_english / english_bengali
chinese_simplified_english / english_chinese_simplified
english_english
french_english /… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/speech-translation-and-summarization.simsamu
Simsamu dataset
This repository contains recordings of simulated medical dispatch dialogs in the
french language, annotated for diarization and transcription. It is published
under the MIT license.
These dialogs were recorded as part of the training of emergency medicine
interns, which consisted in simulating a medical dispatch call where the interns
took turns playing the caller and the regulating doctor.
Each situation was decided randomly in advance, blind to who was playing the… See the full description on the dataset page: https://huggingface.co/datasets/medkit/simsamu.MSU-Benchmark
MSU-Bench: Towards Speaker-Centric Understanding in Conversational Multi-Speaker Scenarios
Interspeech 2026 · ASLP@NPU (Northwestern Polytechnical University), in collaboration with Li Auto.
Zhaokai Sun*, Shuai Wang*, Zhennan Lin*, Chengyou Wang, Dehui Gao, Yuang Cao, Chunjiang He, Pan Zhou, Lei Xie**
Audio, Speech and Language Processing Group (ASLP@NPU), School of Software, Northwestern Polytechnical University, China
School of Intelligent Science and Technology, Nanjing… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/MSU-Benchmark.ModalityFaultLines-SCEval
SCEval — Modality Fault Lines
Data for Modality Fault Lines: Structural Corruptions Reveal Fragile Omni-Modal Reasoning (Findings of EMNLP 2026).
SCEval is a human-verified benchmark for omni-modal robustness. Text, vision, and audio all remain
present, but controlled corruptions make the evidence inside a channel unreliable. Each corrupted
item is paired with its clean counterpart at the example level, so clean-to-corrupted comparisons are
made on the same underlying question… See the full description on the dataset page: https://huggingface.co/datasets/KZL96/ModalityFaultLines-SCEval.MMOU
MMOU
Massive Multi-Task Omni Understanding and Reasoning
Benchmark for Long and Complex Real-World Videos
Project Page
·
HuggingFace
·
Videos (Community Hosted)
·
Paper
·
Evaluator
MMOU evaluates joint audio-visual understanding and reasoning in long and complex real-world videos.
Dataset Summary
MMOU is a benchmark for evaluating whether multimodal models can jointly reason over video, speech, sound, music, and long-range temporal context in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MMOU.MMAG
MMAG: A Multi‑Control Mixed Audio Generation Benchmark
MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
Dataset Structure
The dataset is organized into three subsets, each targeting a specific… See the full description on the dataset page: https://huggingface.co/datasets/rookie9/MMAG.music2chords_v2nano4m-Audio
nano4M-Audio — Team (week-1)
Week-1 data preparation for nano4M-Audio, an extension of EPFL's
nano4M (the educational nano version of
4M / 4M-21)
that adds audio as a fifth modality alongside RGB, depth, surface normals
and captions.
This dataset covers all 12 VGGSound classes assigned to the three-person team:
person
classes
1 (Hassan)
lions roaring, horse neighing, pig oinking, cow lowing
2 (Ziyad)
dog barking, cat meowing, coyote howling, elephant trumpeting
3… See the full description on the dataset page: https://huggingface.co/datasets/zed-m97/nano4m-Audio.OpenGameArt-Mixed-Licenses
Dataset Card for OpenGameArt-Mixed-Licenses
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.AVE-Speech
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
Abstract
AVE Speech is a large-scale Mandarin speech corpus that pairs synchronized audio, lip video and surface electromyography (EMG) recordings. The dataset contains 100 sentences read by 100 native speakers. Each participant repeated the full corpus ten times, yielding over 55 hours of data per modality. These complementary signals enable… See the full description on the dataset page: https://huggingface.co/datasets/MML-Group/AVE-Speech.MF-Skills🚨 Please request access with your institutional email to get access to the dataset.
MF-Skills Dataset
Project page | Paper | Code
Dataset Description
MF-Skills is a large-scale dataset for advancing expert-level music understanding and reasoning in (large) audio-language models. It builds upon audio samples from LAION-DISCO and augments them with rich metadata extracted using a suite of open-source large audio-language models (LALMs) and specialized music analysis… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/MF-Skills.multichannel-meetings-10h
GroundTruth Multi-Channel Meeting Audio Dataset (10h)
Summary
This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant.
Each meeting includes:
One full meeting recording (room microphone)
Individual close-talk recordings for each participant (one file per speaker)
Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.MixEval-X
🚀 Project Page | 📜 arXiv | 👨💻 Github | 🏆 Leaderboard | 📝 blog | 🤗 HF Paper | 𝕏 Twitter
MixEval-X encompasses eight input-output modality combinations and can be further extended. Its data points reflect real-world task distributions. The last grid presents the scores of frontier organizations’ flagship models on MixEval-X, normalized to a 0-100 scale, with MMG tasks using win rates instead of Elo. Section C of the paper presents example data samples and model responses.… See the full description on the dataset page: https://huggingface.co/datasets/MixEval/MixEval-X.cs-envi-dual-encoder-60MM-ContextASR-Bench
MM-ContextASR Bench
Metadata and evaluation splits for Multimodal Conversational Context for
LLM-Based ASR: Data Construction, Training, and Benchmark.
Dataset summary
Config
Examples
Audio
Context
Primary metric
mm_contextasr
1,250 (250 current utterances × 5 histories)
1,439 WAV files included
Controlled user-assistant dialogue
entity Recall
kespeech
19,212
Source ID only
Same-speaker speech and transcript
CER, SER, entity Recall
cv_yue
3,525… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/MM-ContextASR-Bench.EMOjam-actions-v1
jam-actions-v1
Schema: jam-actions-v1/1.0.0 · Version: 1.1.0 · Records: 213 (154 train / 59 test, split by song) ·
Songs: 11 · Families: 9 · Licence: CC-BY-SA-3.0-DE ·
Source repo: mcp-tool-shop-org/ai-jam-sessions
The successor to jam-actions-v0.
Where v0 asked whether a model could use the tools, v1 asks whether a small model can reason from
what the tools return — and it exists in its current shape because, seven training runs in a row,
the answer depended on what the… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1.hillenbrand_vowels
Hillenbrand Vowel Dataset
The Hillenbrand dataset contains recordings of American English vowels produced by speakers from four demographic groups: men, women, boys, and girls.Each audio sample is accompanied by frame-level formant tracks (F1, F2, F3, F4) extracted every 10 ms.
This repository provides the dataset in a structure compatible with the Hugging Face datasets library for easy loading and processing.
Dataset Summary
Sampling rate: 16 kHz
Frame shift: 10… See the full description on the dataset page: https://huggingface.co/datasets/MLSpeech/hillenbrand_vowels.MutiEmo-Test
MultiEmo-Test
MultiEmo-Test is an English evaluation set for instruction-following multi-emotion text-to-speech synthesis. It accompanies HybridEmo, a system for modeling sequential emotion trajectories and simultaneous emotion blending within an utterance.
The dataset is intended for evaluation only. It contains synthesis text, natural-language emotion instructions, emotion annotations, and prompt audio for speaker-timbre conditioning. It does not contain target synthesized… See the full description on the dataset page: https://huggingface.co/datasets/ICTNLP/MutiEmo-Test.instructtts-three-model-gemini-zh
InstructTTSEval 三模型 Gemini 评测数据
本目录整理了 InstructTTSEval 中文集上三个 TTS 模型的生成音频和 Gemini 一致性评测结果:Qwen3-TTS-12Hz-1.7B-VoiceDesign、Seed-Audio-1.0、VoxCPM2。
字段
records.jsonl 每行对应一个模型和一种控制格式(APS、DSD 或 RP):
id:InstructTTSEval 样本 ID
mode:控制格式
model、model_name:模型标识
text:合成文本
instruction:历史评测记录中的输入控制指令,按本行 APS/DSD/RP 格式保留;不等同于各模型 API 的完整请求封装
generated_audio:该模型生成音频的相对路径
reference_audio:原始参考音频的相对路径
gemini_consistent:Gemini judge 的一致性判断
inconsistency_reason:判断为不一致时的原因… See the full description on the dataset page: https://huggingface.co/datasets/zsy814/instructtts-three-model-gemini-zh.jam-actions-v1-probe
jam-actions-v1-probe
Schema: jam-actions-v1-probe/1.0.0 · Records: 24, all split: test · Evaluation only ·
Companion to: jam-actions-v1
Why it exists
An adapter trained on an earlier version of the corpus scored 47/54 on held-out acoustic takes.
Its completions, which state the comparison before the label, showed that it wrote against a 50-cent gate whenever it saw a minus sign — and negative cents occurred in exactly one class of that
corpus. The main split could… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v1-probe.EGYSpeak
EGYSpeak
A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline.
Quick Start
1. Download the dataset:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="MohamedGomaa30/EGYSpeak",
repo_type="dataset",
local_dir="EGYSpeak",
)
2. Extract the dataset:
from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.Music
Made by herza For APP.t
jam-actions-acoustic-v0
Dataset Card for jam-actions-acoustic-v0
Version: 1.0.2
Published at mcp-tool-shop/jam-actions-acoustic-v0. No DOI.
Summary
108 constructible gold records of grounded MCP tool use over monophonic audio analysis. Each record pairs a 4-note right-hand reduction of a public-domain library phrase with a seeded synthetic take and a gold verdict (match, pitch fail/warn, timing fail/pass, missed, extra, in-tune vibrato, or nothing-to-grade silence).
This is not a musical… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-acoustic-v0.my-voxtral-dataset
