datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audioset
AudioSet
AudioSet[1] consists of an expanding ontology of 527 audio event classes and a collection of 2M human-labelled 10-second sound clips drawn from YouTube.
Some clips are missing on YouTube, so the number of files downloaded is different from time to time.
This repository contains 20550 / 22160 of the balanced train set, 1913637 / 2041789 of the unbalanced train set (separated into 41 parts), and 18887 / 20371 of the evaluation set.
The pre-process script can be found at… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/audioset.Neko_Audio-80K_Short
Audio2Tool
Audio2Tool: Speak, Call, Act — A Dataset for Benchmarking Speech Tool Use
Authors: Ramit Pahwa1,∗,∗∗, Apoorva Beedu1,∗, Parivesh Priye1, Rutu Gandhi†1, Saloni Takawale†1, Aruna Baijal1, Zengli Yang1
1 Rivian & Volkswagen Technologies · ∗ equal contribution · ∗∗ corresponding author · † equal contribution
📄 Project page / demo: https://audio2tool.github.io/
📦 Dataset: https://huggingface.co/datasets/RVtech/Audio2Tool
✉️ Contact (corresponding… See the full description on the dataset page: https://huggingface.co/datasets/RVtech/Audio2Tool.Neko_Audio-30K_LongAudioMCQ-StrongAC-GeminiCoT
[ICLR 2026] [DCASE 2026 Training Set] AudioMCQ-StrongAC-GeminiCoT
This dataset is a highly curated subset of the AudioMCQ dataset, containing samples with native Chain-of-Thought (CoT) reasoning from Gemini 3.1 Pro that were answered correctly.
Note: The native CoT reasoning has been summarized by Gemini's internal algorithm before output, yet still retains rich audio details including timestamps, acoustic descriptions, and step-by-step temporal analysis.
🏆 DCASE 2026… See the full description on the dataset page: https://huggingface.co/datasets/Harland/AudioMCQ-StrongAC-GeminiCoT.AudioJailbreak
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly.
📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.ASID-1M
ASID-1M: Attribute-Structured and Quality-Verified Audiovisual Instructions
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce ASID-1M, a large-scale audiovisual instruction dataset built to support universal video understanding with fine-grained, controllable supervision.
Most existing video-instruction data represents complex audiovisual content as a single, monolithic caption. This often leads to incomplete coverage (missing audio… See the full description on the dataset page: https://huggingface.co/datasets/AudioVisual-Caption/ASID-1M.raw_tts_esc_ESPnet_espnet_mls-audioset_soundstream_16kAudioMCQ
[ICLR 2026] AudioMCQ: Audio Multiple-Choice Question Dataset
Also the official repository for the paper "Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models"
News
[2026.04] Update on MMSU Metric of released models: Based on community feedback, we identified a flaw in our evaluation script that artificially inflated the MMSU scores of our released models by ignoring sequence order. We sincerely apologize for… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/AudioMCQ.nano4m-Audio
nano4M-Audio — Team (week-1)
Week-1 data preparation for nano4M-Audio, an extension of EPFL's
nano4M (the educational nano version of
4M / 4M-21)
that adds audio as a fifth modality alongside RGB, depth, surface normals
and captions.
This dataset covers all 12 VGGSound classes assigned to the three-person team:
person
classes
1 (Hassan)
lions roaring, horse neighing, pig oinking, cow lowing
2 (Ziyad)
dog barking, cat meowing, coyote howling, elephant trumpeting
3… See the full description on the dataset page: https://huggingface.co/datasets/zed-m97/nano4m-Audio.bengali-talkshow-audio
Bengali Talkshow Audio Dataset
A large-scale collection of 1,180 Bengali talk show audio recordings totaling 789+ hours of multi-speaker speech, sourced from Bangladeshi television talk shows and political debate programs.
Dataset Description
This dataset contains audio from Bengali-language TV talk shows, political debates, and news discussion programs from major Bangladeshi television channels. Each recording features multiple speakers engaged in discussion, making it… See the full description on the dataset page: https://huggingface.co/datasets/nymtheescobar/bengali-talkshow-audio.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.homorich-negara-gooya-pron-regen-audioaudio
CoDaCo - audios dataset
This dataset was created using codaco.app.
Description
All data contributed to this campaign goes to the global CoDaCo datasets.
Labels
This dataset includes the following labels:
Spoken text
Tags
Emotions
AI generated
Quality rating
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any purpose, including commercially, as long as
you give appropriate credit. See LICENSE… See the full description on the dataset page: https://huggingface.co/datasets/codaco/audio.audio-event-classification-post-public
audio-event-classification-post-public
Sound-event and acoustic-scene classification annotations: ESC-50 (environmental), UrbanSound8K, FSD50k (50k+ events), TUT-Acoustic-Scenes-2017, DCASE-2025, NonSpeech7k (vocal sounds), VocalSound (laugh/cough/sigh). Useful for training audio LLMs on the perception substrate underneath higher-level reasoning.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-event-classification-post-public.Selena-Gomez-With-Lyrics-And-Spotify-Audio-Featuresorca-audio-qa-annotations
ORCA Audio QA Annotations
Annotation data for training and evaluating ORCA (Open-ended Response Correctness Assessment), a scoring model for audio question-answering tasks.
Paper: ORCA: Open-ended Response Correctness Assessment for Audio Question Answering — accepted to TACL 2026
Code & usage: github.com/BUTSpeechFIT/ORCA
Pretrained Models:
orca-olmo-2-1b-multinomial
orca-gemma-3-4b-it-multinomial
orca-llama-3.2-3b-it-multinomial
Dataset overview
ORCA is… See the full description on the dataset page: https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations.audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.PICSAFEv1
Speech Quality Test Labels
It publishes only this README and metadata.jsonl, because the 14 source datasets have different copyright and access terms.
The manifest contains the 10,728-sample PICSAFEv1 subset. It is not a license to redistribute the underlying recordings.
Metadata fields
file_name: audio basename.
id: globally unique sample identifier with the [source] prefix.
source: source dataset name.
source_path: path relative to the origin directory.
tags:… See the full description on the dataset page: https://huggingface.co/datasets/AudioCC-Lab/PICSAFEv1.AVQA-Audio-Rubrics
AVQA Audio-Reasoning Rubrics
Project Page | Paper | Code
Audio-grounded, binary-evaluable evaluation rubrics for the full
AVQA training set, generated for
process-level reward modeling in audio reasoning RL (e.g. GRPO / RLHF with
rubric-as-reward).
Each training question is annotated with 5 rubrics, one per evaluation
facet, that judge the quality of an audio-reasoning response — not just final
answer correctness. The rubrics are designed to be scored Yes/No by an
LLM judge that… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/AVQA-Audio-Rubrics.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.audio-event-triage-20260823-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260823-dataset.audio-event-triage-20260902-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260902-dataset.AudioSkills-Llama3
AudioSkills-XL Dataset
To promote the development of open source models, we have released AudioSkills using the exact same method generated with Llama 3.1-8B Instruct instead of GPT4o in the original.
Project page | Paper | Code
Dataset Description
AudioSkills-XL is a large-scale audio question-answering (AQA) dataset designed to develop (large) audio-language models on expert-level reasoning and problem-solving tasks over short audio clips (≤30 seconds). It… See the full description on the dataset page: https://huggingface.co/datasets/sonalkum/AudioSkills-Llama3.Audio-Cogito
Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models
Audio-Cogito is a large-scale audio reasoning dataset introduced in the paper Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models.
The released data contains 545k high-quality audio reasoning samples spanning sound, speech, and music domains. Each sample includes label annotations, Chain-of-Thought (CoT) annotations, and final answers.
Links… See the full description on the dataset page: https://huggingface.co/datasets/lilonghao/Audio-Cogito.seamless-bg
Seamless Background Robustness Pilot
V3 expansion available: v3/README.md documents the expanded 1,985-event pool. Use v3/events_all.jsonl and v3/clips_all.jsonl for combined manifests. The original pilot statistics and files below remain unchanged.
A compact, paired-audio candidate pool for incremental full-duplex interaction alignment and later background-speech augmentation. Derived from Meta's Seamless Interaction, by selecting events from the original train split only. This… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/seamless-bg.indicvoices-r-audiogod_level_beat_producer_big_fish_audio
God-Level Beat Producer — Big Fish Audio + All Major DAWs
The ultimate dataset for training an elite AI music producer
This dataset trains LLMs to become God-Level Beat Producers specializing in:
Big Fish Audio sample packs & loops
Professional beat making across all genres (Trap, Drill, Melodic, Lo-Fi, House, Afrobeats, etc.)
Complete DAW workflows (FL Studio, Logic Pro, Ableton Live, Reason, Cakewalk, Cubase, Pro Tools)
Virtual instruments (Serum, Vital, Omnisphere, Kontakt… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_beat_producer_big_fish_audio.sf-streets-gpt-audio-resultsaudio-event-triage-20260912-dataset
Audio Event Triage Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Operations teams need an explainable starting point for classifying alarms, machinery noise, and speech-like events.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/audio-event-triage-20260912-dataset.
