datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dcase2025-audio-qa
DCASE 2025 HuggingFace Dataset
This script creates a HuggingFace dataset from the DCASE 2025 Audio Question Answering data.
Dataset Structure
The dataset contains the following columns:
audio: Audio file (automatically converted to mono 16bit 48kHz)
question: The formatted question with choices (if applicable)
question_text: The original question text without choices
answer: The correct answer
id: Unique identifier for each example
audio_url: Original audio URL from the… See the full description on the dataset page: https://huggingface.co/datasets/gijs/dcase2025-audio-qa.DCASE2026-Task5-DevSet
DCASE 2026 Task 5 Audio-Dependent Question Answering (ADQA) Development Set
This is the official Development Set for DCASE 2026 Challenge Task 5: Audio-Dependent Question Answering (ADQA).
The ADQA task focuses on addressing "Textual Hallucination" in Large Audio-Language Models (LALMs) — where models pass audio understanding benchmarks by relying on text prompts and internal linguistic priors rather than actual audio perception. ADQA introduces a rigorous evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Harland/DCASE2026-Task5-DevSet.2025_DCASE_AudioQA_Official
Audio SFT / Post-Training Data
The proposed audio question answering (AQA) dataset
with three categories: Bioacoustics QA (BQA), Temporal Soundscapes QA (TSQA), and Complex QA (CQA)
DCASE 2025 Task Description
Audio QA Model Baseline
Watkins Marine Mammal Sound Database
"Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
📢 Post-Challenge Research Note
While the DCASE 2025 Challenge… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official.dcase2025_task2_dev
DCASE 2025 Task 2 - Development Dataset
Attributes d1v, d2v, d3v are encoded as ClassLabels.
Usage
from datasets import load_dataset
dataset = load_dataset('HTill/dcase2025_task2_dev', trust_remote_code=True)
dcase2016_task2_extract_unit
Dataset Card for "dcase2016_task2_extract_unit"
More Information needed
dcase2016_task2_synth
Dataset Card for "dcase2016_task2_synth"
More Information needed
dcase2025-audio-qa
DCASE 2025 HuggingFace Dataset
This script creates a HuggingFace dataset from the DCASE 2025 Audio Question Answering data.
Dataset Structure
The dataset contains the following columns:
audio: Audio file (automatically converted to mono 16bit 48kHz)
question: The formatted question with choices (if applicable)
question_text: The original question text without choices
answer: The correct answer
id: Unique identifier for each example
audio_url: Original audio URL… See the full description on the dataset page: https://huggingface.co/datasets/xppp1983/dcase2025-audio-qa.beans_dcase
Dataset Card for "beans_dcase"
Dataset Description
Paper: https://dcase.community/documents/workshop2021/proceedings/DCASE2021Workshop_Morfi_52.pdf
Dataset Summary
This is the dataset used for the DCASE 2021 Task and contains annotated mammal and bird multi-species recordings with splits and preprocessing like described in BEANS. It is used for detection tasks.
Data Splits
train
train_low
valid
test
702
151
234
232
dcase24_task10_loc1
DCASE 2024 Challenge Task 10 Development Dataset: Acoustic-based Traffic Monitoring - Location 1 subset
Citation
Bondi, L., Ghaffarzadegan, S., Damiano, S., Kumar, A., Wu, H.-H., Lin, W.-C., Das, S., Horst, H.-G., & Waterschoot, T. van . (2024). DCASE 2024 Challenge Task 10 Development Dataset: Acoustic-based Traffic Monitoring [Data set]. Zenodo. https://doi.org/10.5281/zenodo.10700792
License
Creative Commons Attribution-NonCommercial-ShareAlike 4.0… See the full description on the dataset page: https://huggingface.co/datasets/renumics/dcase24_task10_loc1.AudioMomentRetrievalFromLongAudio_DCASE2026EvaluationData
What is this?
This repository contains data for DCASE Challenge 2026 Task 6.
Audio and text features using CLAP and a sliding window, following the same feature extraction protocol as the CASTELLA dataset.
submission template
File structure
clap
└──dcase2026_evaluation_audio_{vid}.npz
clap_text
└──qiddcase2026_evaluation_q{qid}.npz
Raw audio files
If participants require the raw audio, please contact the… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/AudioMomentRetrievalFromLongAudio_DCASE2026EvaluationData.dcase2025HEARSoundEventDetection_DCASE2016Task2AudibleLight_Eigenmike32-5_DCASE-STARSS23_Dataset
AudibleLight Eigenmike32-5 DCASE-STARSS23 Dataset
This dataset contains 121 synthetic spatial audio scenes — 111 for training and 10 for evaluation — of 60 seconds each, generated with the AudibleLight dataset generator (DOI). Each scene is rendered as five independent simulated Eigenmike32 captures, with 32 channels per capture, resulting in 570 minutes of multichannel audio in total at 24 kHz.
Foreground Audio
Foreground events are sampled from ESC-50: Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/PhilippXXY/AudibleLight_Eigenmike32-5_DCASE-STARSS23_Dataset.DCase2016_Task2_Hear_2021DCASE_4products_and_marketing_emails
Dataset Card for "products_and_marketing_emails"
More Information needed
DCASE_2025_AudioQA_backupDCASE 2025 Audio Question and Answering
2025_DCASE_AudioQA
