datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dcase2025-audio-qa
DCASE 2025 HuggingFace Dataset
This script creates a HuggingFace dataset from the DCASE 2025 Audio Question Answering data.
Dataset Structure
The dataset contains the following columns:
audio: Audio file (automatically converted to mono 16bit 48kHz)
question: The formatted question with choices (if applicable)
question_text: The original question text without choices
answer: The correct answer
id: Unique identifier for each example
audio_url: Original audio URL from the… See the full description on the dataset page: https://huggingface.co/datasets/gijs/dcase2025-audio-qa.AudioQA-1M2025_DCASE_AudioQA_Official
Audio SFT / Post-Training Data
The proposed audio question answering (AQA) dataset
with three categories: Bioacoustics QA (BQA), Temporal Soundscapes QA (TSQA), and Complex QA (CQA)
DCASE 2025 Task Description
Audio QA Model Baseline
Watkins Marine Mammal Sound Database
"Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
📢 Post-Challenge Research Note
While the DCASE 2025 Challenge… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official.AudioQA-1Mqaida-audioorca-audio-qa-annotations
ORCA Audio QA Annotations
Annotation data for training and evaluating ORCA (Open-ended Response Correctness Assessment), a scoring model for audio question-answering tasks.
Paper: ORCA: Open-ended Response Correctness Assessment for Audio Question Answering — accepted to TACL 2026
Code & usage: github.com/BUTSpeechFIT/ORCA
Pretrained Models:
orca-olmo-2-1b-multinomial
orca-gemma-3-4b-it-multinomial
orca-llama-3.2-3b-it-multinomial
Dataset overview
ORCA is… See the full description on the dataset page: https://huggingface.co/datasets/BUT-FIT/orca-audio-qa-annotations.nlp-qa-audio-video
NLP QA Audio Video Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for NLP QA work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/leond-u0114/nlp-qa-audio-video.dcase2025-audio-qa
DCASE 2025 HuggingFace Dataset
This script creates a HuggingFace dataset from the DCASE 2025 Audio Question Answering data.
Dataset Structure
The dataset contains the following columns:
audio: Audio file (automatically converted to mono 16bit 48kHz)
question: The formatted question with choices (if applicable)
question_text: The original question text without choices
answer: The correct answer
id: Unique identifier for each example
audio_url: Original audio URL… See the full description on the dataset page: https://huggingface.co/datasets/xppp1983/dcase2025-audio-qa.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.nlp-qa-audio-video
NLP QA Audio Video Data Notes
Dataset summary
This data card accompanies a lightweight NLP QA loader for Audio Video metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/ShwetaMsm3867/nlp-qa-audio-video.nlp-qa-audio-video
NLP QA Audio Video Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for NLP QA work with Audio Video inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/lucasjlpt/nlp-qa-audio-video.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.trivia_qa-audiotrivia_qa-audioAudio_QA_datasetnlp-qa-image-audio
NLP QA Image Audio Data Notes
Dataset summary
This data card accompanies a lightweight NLP QA loader for Image Audio metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/shwetasharmabury/nlp-qa-image-audio.sample_qa_audio_datasetaudio_replay-14_trivia_qa-audioaudio_no-replay-14_trivia_qa-audiotext_replay-14_trivia_qa-audiotrivia_qa-audio-scoreaudio_L2-regular-14_trivia_qa-audiodataset_127256553_nlp_qa_audio_text
dataset_127256553_nlp_qa_audio_text.py
Dataset Summary
A nlp qa dataset with audio text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: domain specific
Augmentation: light
Splits & Sampling
Split strategy: random 90 10
Sampling: hard negative
Quality & Labeling
Quality filtering: strict
Labeling: self training
Files
dataset_127256553_nlp_qa_audio_text.py — main… See the full description on the dataset page: https://huggingface.co/datasets/Iusokolov/dataset_127256553_nlp_qa_audio_text.audio_merge-ties_trivia_qa-audioaudio_L2-regular-linear_trivia_qa-audioaudio_L2-regular-dare_trivia_qa-audioaudio_L2-regular-ties_trivia_qa-audioaudio_tts_qatar_wavgemini-audio-qa-pairstrivia_qa-audio-text
