datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multispeech_qa
MultispeechQA Dataset
Multilingual spoken-audio question-answering dataset covering 16 languages: Arabic, Czech, German, Greek, French, Hebrew, Hindi, Indonesian, Japanese, Korean, Dutch, Portuguese, Romanian, Spanish, Turkish, Ukrainian. Each example pairs an audio clip with a question and its answer.
Each language config has train / test / validation splits, sharded as multiple .parquet files.
from datasets import load_dataset
ds = load_dataset("your-username/multispeech-qa"… See the full description on the dataset page: https://huggingface.co/datasets/tolulope/multispeech_qa.muhsin-al-qasim-192kbps
Muhsin Al-Qasim
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
muhsin-al-qasim-192kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
192 kbps
Ayah files
6350 (1690 MiB)
Verified against the upstream MD5 list
6350
Ayahs absent upstream
0
Upstream folder
Muhsin_Al_Qasim_192kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3 is… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/muhsin-al-qasim-192kbps.khalid-al-qahtani-192kbps
Khalid Abdullah Al-Qahtani
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
khalid-al-qahtani-192kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
192 kbps
Ayah files
6350 (2227 MiB)
Verified against the upstream MD5 list
6350
Ayahs absent upstream
0
Upstream folder
Khaalid_Abdullaah_al-Qahtaanee_192kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/khalid-al-qahtani-192kbps.MECAT-QAMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-Caption (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-QA.dcase2025-audio-qa
DCASE 2025 HuggingFace Dataset
This script creates a HuggingFace dataset from the DCASE 2025 Audio Question Answering data.
Dataset Structure
The dataset contains the following columns:
audio: Audio file (automatically converted to mono 16bit 48kHz)
question: The formatted question with choices (if applicable)
question_text: The original question text without choices
answer: The correct answer
id: Unique identifier for each example
audio_url: Original audio URL from the… See the full description on the dataset page: https://huggingface.co/datasets/gijs/dcase2025-audio-qa.preprocessed_speech_datasetschinese_qaVox-Infinity-QAaudioset-20k-qwen2.5-7b-qaSD-QAnasser-al-qatami-128kbps
Nasser Al-Qatami
Part of Maqra, an open, verified archive of verse-by-verse Qur'an recitations mirrored from everyayah.com.
Set
nasser-al-qatami-128kbps
Style
murattal
Riwayah
hafs
Kind
recitation
Bitrate
128 kbps
Ayah files
6350 (1126 MiB)
Verified against the upstream MD5 list
6350
Ayahs absent upstream
0
Upstream folder
Nasser_Alqatami_128kbps
Files
One MP3 per ayah, named SSSAAA.mp3 (surah 3 digits, ayah 3 digits). 001001.mp3… See the full description on the dataset page: https://huggingface.co/datasets/maqra-project/nasser-al-qatami-128kbps.qaloon_datasetqasr-arabic-speech-continuationsqaida-audiopublic_sg_speech_qa_test@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and Zhang, Wenyu and Liu, Zhengyuan and Aw, AiTi and Chen, Nancy F},
journal={NAACL},
year={2025}
}
NSF-QA
NSF-QA
NSF-QA is a question–answering dataset for meeting understanding, derived from the
NOTSOFAR-1 distant meeting
transcription corpus (CHiME-8 Task 2). The name stands for NOTSOFAR-QA.
Each example pairs a NOTSOFAR-1 meeting (its transcript / speaker-attributed
content) with one or more question–answer items, enabling evaluation of models on
reading comprehension and reasoning over real, multi-speaker meeting conversations.
How it was built
NSF-QA is a… See the full description on the dataset page: https://huggingface.co/datasets/popcornell/NSF-QA.speech-triavia-qa
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset speech-triviaqa --model gpt4o_speech
python audio_evals/main.py --dataset speech-triviaqa-s2t --model gpt4o_speech
🚀超凡体验,尽在UltraEval-Audio🚀
UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效:
一键式基准管理… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-triavia-qa.dcase2025-audio-qa
DCASE 2025 HuggingFace Dataset
This script creates a HuggingFace dataset from the DCASE 2025 Audio Question Answering data.
Dataset Structure
The dataset contains the following columns:
audio: Audio file (automatically converted to mono 16bit 48kHz)
question: The formatted question with choices (if applicable)
question_text: The original question text without choices
answer: The correct answer
id: Unique identifier for each example
audio_url: Original audio URL… See the full description on the dataset page: https://huggingface.co/datasets/xppp1983/dcase2025-audio-qa.audiocaps_qa_test@inproceedings{kim2019audiocaps,
title={Audiocaps: Generating captions for audios in the wild},
author={Kim, Chris Dongjoo and Kim, Byeongchang and Lee, Hyunmin and Kim, Gunhee},
booktitle={Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)},
pages={119--132},
year={2019}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/audiocaps_qa_test.audio-reasoning-qa-post-public
audio-reasoning-qa-post-public
Question-answering and multi-task audio reasoning annotations across 15 public audio QA datasets. Spans general audio QA (Clotho-AQA, HeySQuAD), music reasoning (MU-LLaMA, MusicBench, LLARK-MTAT, Music-AVQA), speech-grounded QA (LibriSQA, GigaSpeech), and NVIDIA-aggregator skill subsets (TemporalQA, CountingQA, AudioSet-Speech-QA, GigaSpeech-Long-QA). Closes a substantial slice of the public audio-reasoning SFT gap (compare to NVIDIA AudioSkills-XL… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-reasoning-qa-post-public.wavcaps_qa_test@article{mei2024wavcaps,
title={Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research},
author={Mei, Xinhao and Meng, Chutong and Liu, Haohe and Kong, Qiuqiang and Ko, Tom and Zhao, Chengqi and Plumbley, Mark D and Zou, Yuexian and Wang, Wenwu},
journal={IEEE/ACM Transactions on Audio, Speech, and Language Processing},
year={2024},
publisher={IEEE}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/wavcaps_qa_test.NigBench-MAMAI-Speech-QA
Voices for Smart Care
A Rural-First Multilingual Voice Dataset for Maternal Health in Nigeria
Voices for Smart Care is a multilingual speech dataset containing real-world maternal and reproductive health questions collected from women across Nigeria. The dataset was created to support the development and evaluation of Automatic Speech Recognition (ASR) and Large Language Models (LLMs) for low-resource African languages in healthcare settings.
Unlike generic speech… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/NigBench-MAMAI-Speech-QA.qa_BBQ_trans_gender
Dataset Card for "qa_BBQ_trans_gender"
More Information needed
quranevalMonerProject_qalonuniverse_qa_check_dup
UniVerse: BENCHMARKING AND ENHANCING LALMS ON LOW-RESOURCE FOLK MUSIC
qa_BBQ_bi_gender
Dataset Card for "qa_BBQ_bi_gender"
More Information needed
hausa-qa-1k
hausa-qa-1k
Hausa telecom question-answer pairs for conversational voice agents
(Ultravox-style: spoken customer query + agent response text).
Question-answer pairs for training a conversational AI voice agent in the telecommunications industry.
Topics: billing_and_payments, plans_and_pricing, network_and_connectivity, data_and_usage, device_and_sim, customer_account, international_services, technical_support, contracts_and_policies, value_added_services, fiber_and_broadband… See the full description on the dataset page: https://huggingface.co/datasets/vaghawan/hausa-qa-1k.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.trivia_qa-audiospeech-qa-magahi-hi-karya
