audio-qa
Datasets
All datasets matching “audio-qa”dcase2025-audio-qa
DCASE 2025 HuggingFace Dataset
This script creates a HuggingFace dataset from the DCASE 2025 Audio Question Answering data.
Dataset Structure
The dataset contains the following columns:
audio: Audio file (automatically converted to mono 16bit 48kHz)
question: The formatted question with choices (if applicable)
question_text: The original question text without choices
answer: The correct answer
id: Unique identifier for each example
audio_url: Original audio URL from the… See the full description on the dataset page: https://huggingface.co/datasets/gijs/dcase2025-audio-qa.AudioQA-1M2025_DCASE_AudioQA_Official
Audio SFT / Post-Training Data
The proposed audio question answering (AQA) dataset
with three categories: Bioacoustics QA (BQA), Temporal Soundscapes QA (TSQA), and Complex QA (CQA)
DCASE 2025 Task Description
Audio QA Model Baseline
Watkins Marine Mammal Sound Database
"Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
📢 Post-Challenge Research Note
While the DCASE 2025 Challenge… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official.AudioQA-1Mqaida-audioorca-audio-qa-annotations
ORCA Audio QA Annotations
Annotation data for training and evaluating ORCA (Open-ended Response Correctness Assessment), a scoring model for audio question-answering tasks.
Paper: ORCA: Open-ended Response Correctness Assessment for Audio Question Answering — accepted to TACL 2026
Code & usage: github.com/BUTSpeechFIT/ORCA
Pretrained Models:
orca-olmo-2-1b-multinomial
orca-gemma-3-4b-it-multinomial
orca-llama-3.2-3b-it-multinomial
Dataset overview
ORCA is… See the full description on the dataset page: https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations.
