datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AudioMarathon
🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs
Abstract
AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars:
long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.AudioJailbreak
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly.
📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.sea_audiobench_datasets_TCQ
SEA-SpeechBench — TCQ (Timestamped Content Query)
This dataset is the timestamped-content-query (TCQ) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 14,172 evaluation examples
across five languages, each pairing an audio recording with an instruction
and a reference answer.
Given the recording and a timestamp, a model must report what is said at
that point in the audio. Contexts run from 30 seconds to 3… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TCQ.python-audio-copilot-training-using-function-knowledge-graphs
Python Copilot Audio Training using Global Functions with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each global function has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-function-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs.python-audio-copilot-training-using-import-knowledge-graphs
Python Copilot Audio Training using Imports with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each imported module for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-import-knowledge-graphs.python-audio-copilot-training-using-inheritance-knowledge-graphs
Python Copilot Audio Training using Inheritance and Polymorphism Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each base class for each unique class in each module file has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-inheritance-knowledge-graphs.python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27
Python Copilot Audio Training using Class with Knowledge Graphs
This dataset is a subset of the matlok python copilot datasets. Please refer to the Multimodal Python Copilot Training Overview for more details on how to use this dataset.
Details
Each class method has a question and answer mp3 where one voice reads the question and another voice reads the answer. Both mp3s are stored in the parquet dbytes column and the associated source code file_path identifier.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/matlok/python-audio-copilot-training-using-class-knowledge-graphs-2024-01-27.sea_audiobench_datasets_TLoc
SEA-SpeechBench — TLoc (Temporal Localization)
This dataset is the temporal-localization (TLoc) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 6,736 evaluation examples
across five languages, each pairing an audio recording of 30–180 seconds
with an instruction and a reference answer.
Given the recording and the instruction, a model must identify when a
described utterance occurs.
Quick start… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TLoc.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.coqar-clarifications-audio
CoQAR Clarifications with synthetic context audio
These audio recordings are AI-generated speech, not recordings of human speakers.
OpenAI tts-1 narrated each exact story using voice alloy, speed 1,
and MP3 output. Long stories are synthesized in ordered parts and joined; see the
audio generation manifest for part boundaries and measured audio properties.
No questions, answers, rationales, or stored model prompts were narrated.
The original appended and inserted configurations… See the full description on the dataset page: https://huggingface.co/datasets/rvashurin/coqar-clarifications-audio.AudioMarathon
AudioMarathon
AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings. The release package in this directory is organized around 11 benchmark tasks spanning meeting summarization, automatic speech recognition, reading comprehension, authenticity detection, music genre classification, acoustic scene classification, emotion recognition, spoken named entity reasoning, sound event detection, speaker gender… See the full description on the dataset page: https://huggingface.co/datasets/AudioMarathon/AudioMarathon.sea_audiobench_datasets_SQA
SEA-SpeechBench — SQA (Spoken Question Answering)
This dataset is the spoken-question-answering (SQA) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 5,462 evaluation examples
across five languages, each pairing an audio recording with a question and a
reference answer.
Given the recording and the question, a model must answer using the content
of the speech.
Quick start
Requires datasets>=4.0… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_SQA.AdvBench-Audio
AdvBench-Audio
AdvBench-Audio is an audio-version benchmark used in our study to evaluate the safety, jailbreak, and adversarial robustness of audio-language models (ALMs). It contains audio renderings of adversarial prompts paired with their corresponding targets for safety evaluation.
This dataset is used in our paper: ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models.
Dataset at a glance
Format: JSONL + WAV files
Fields… See the full description on the dataset page: https://huggingface.co/datasets/WeifeiJin/AdvBench-Audio.Ficbook-Audio-Instruct-10K
Ficbook Audio Instruct 10K
Synthetic audio instruction dataset for training Russian audio-language models.
Contains ~10K samples of fiction text voiced with OpenAI TTS and paired with diverse instruction tasks.
Dataset Description
This dataset was created for training and evaluating audio-language models on Russian fiction content.
Each sample contains:
Audio: Fiction text voiced using OpenAI's gpt-4o-mini-tts model
Text: Original text from ficbook stories
Question:… See the full description on the dataset page: https://huggingface.co/datasets/Vikhrmodels/Ficbook-Audio-Instruct-10K.SRQA_Audio
SRQA Audio
SRQA Audio is the public audio asset bundle for the synthetic Spoken Reasoning Question Answering (SRQA) benchmark used in the paper Learning When to Think While Listening in Large Audio-Language Models. It provides audio files for evaluating audio-input models on spoken versions of established reasoning tasks.
Contents
Rewritten and TTS-rendered benchmark audio
The following benchmark tracks were rewritten into spoken queries and rendered with the… See the full description on the dataset page: https://huggingface.co/datasets/Oulasong/SRQA_Audio.AVQA-Audio-Rubrics
AVQA Audio-Reasoning Rubrics
Project Page | Paper | Code
Audio-grounded, binary-evaluable evaluation rubrics for the full
AVQA training set, generated for
process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with
rubric-as-reward).
Each training question is annotated with 5 rubrics, one per evaluation
facet, that judge the quality of an audio-reasoning response — not just final
answer correctness. The rubrics are designed to be scored Yes/No by an
LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.Real_Audio_Bench
Real Audio Bench
Real Audio Bench is a compact human-recorded benchmark for streaming spoken reasoning. The benchmark contains 186 English audio recordings. Each item is a natural spoken question in which an early answer can be misleading until a later cue is heard.
Audio files are hosted at: https://huggingface.co/datasets/Oulasong/Real_Audio_Bench
Files
audio/: 186 WAV files.
real_audio_bench_manifest.jsonl: one metadata row per audio file, including the… See the full description on the dataset page: https://huggingface.co/datasets/Oulasong/Real_Audio_Bench.AdvBench-Audio
AdvBench-Audio
AdvBench-Audio is an audio-version benchmark used in our study to evaluate the safety, jailbreak, and adversarial robustness of audio-language models (ALMs). It contains audio renderings of adversarial prompts paired with their corresponding targets for safety evaluation.
This dataset is used in our paper: ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models.
Dataset at a glance
Format: JSONL + WAV files… See the full description on the dataset page: https://huggingface.co/datasets/leopixelleo/AdvBench-Audio.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.torgo_audio_datasetaudio_noise
Audio Noise Recognition Dataset
Overview
This dataset is designed for evaluating AI models' ability to identify and classify various types of noise in audio recordings. It contains 28 carefully curated audio samples covering diverse noise scenarios, including both stationary and non-stationary noise types.
Dataset Structure
audio_noise/
├── test/
│ ├── audio/
│ │ ├── NS_001.wav
│ │ ├── NS_002.wav
│ │ └── ...
│ │ └── NS_028.wav
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/hzyhpp123/audio_noise.raw_audiogene-speech-audio-instruct
speech-audio-instruct v3
Gate-passed instruction data for speech-audio — published when 50 fresh examples cleared the quality bar
Kind: synthetic
Domain: speech-audio
Records: 139
Created: 2026-06-25T15:25:34+00:00
SHA-256: c388150c1632b511928f0569229ed89fa9d28953fe4c9ae44811f8ed4eaf45d2
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7}
Generated by: Qwen3-4B-Instruct-2507-Q4_K_M.gguf (backend: llama)… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-speech-audio-instruct.audio-agent-bench-suite
Audio Agent Bench Suite
A suite of six multi-turn, multi-domain spoken conversational benchmarks for evaluating voice AI and audio agent systems. Each sub-dataset targets a distinct real-world deployment domain, together covering the core capabilities required of production audio agents: instruction following, knowledge-base grounding, tool/function-call accuracy, long-range conversational memory, and state tracking.
Sub-datasets
Dataset
Domain
Turns
HuggingFace… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/audio-agent-bench-suite.audio-loop
Audio Loop — Long-Audio Reasoning Benchmarks
Two evaluation sets for long-form audio reasoning, packaged as one dataset
with two configs:
Config
Task
Audio source
Items
aae_tts
Argument structure & contradiction
TTS-synthesized argumentative essays
20
iq2_qa
Multi-hop QA over debates
Recorded live debates
13
Load annotations
from datasets import load_dataset
aae = load_dataset("audioloop/audio-loop", "aae_tts", split="test")
iq2 =… See the full description on the dataset page: https://huggingface.co/datasets/audioloop/audio-loop.
