datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIR-Bench-Dataset
AIR-Bench
Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks.
The former consists of 19 tasks with approximately 19k single-choice questions.
The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon).
Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.Seamless_Dummy_Dataset_Fixed
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
sea_audiobench_datasets_TCQ
SEA-SpeechBench — TCQ (Timestamped Content Query)
This dataset is the timestamped-content-query (TCQ) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 14,172 evaluation examples
across five languages, each pairing an audio recording with an instruction
and a reference answer.
Given the recording and a timestamp, a model must report what is said at
that point in the audio. Contexts run from 30 seconds to 3… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TCQ.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
CaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.sea_audiobench_datasets_TLoc
SEA-SpeechBench — TLoc (Temporal Localization)
This dataset is the temporal-localization (TLoc) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 6,736 evaluation examples
across five languages, each pairing an audio recording of 30–180 seconds
with an instruction and a reference answer.
Given the recording and the instruction, a model must identify when a
described utterance occurs.
Quick start… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_TLoc.sea_audiobench_datasets_SQA
SEA-SpeechBench — SQA (Spoken Question Answering)
This dataset is the spoken-question-answering (SQA) task of
SEA-SpeechBench, a large-scale multitask benchmark for speech
understanding across Southeast Asia. It contains 5,462 evaluation examples
across five languages, each pairing an audio recording with a question and a
reference answer.
Given the recording and the question, a model must answer using the content
of the speech.
Quick start
Requires datasets>=4.0… See the full description on the dataset page: https://huggingface.co/datasets/MERaLiON/sea_audiobench_datasets_SQA.medmosaic-dataset
MedMosaic Dataset
A comprehensive medical audio question-answering dataset designed for evaluating audio understanding models in clinical and healthcare contexts.
Dataset Description
This dataset contains audio recordings paired with clinical questions and answers across multiple QA types. It is designed to benchmark audio-language models on medical reasoning tasks.
Dataset Structure
The dataset is organized into 7 subfolders, each representing a different QA… See the full description on the dataset page: https://huggingface.co/datasets/icml-anon-submission/medmosaic-dataset.Deaftest_datasetOfficial Deaftest dataset for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?".
🌟 For more details, please refer to the project page with data examples: https://av-odyssey.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface AV-Odyssey Dataset] [🤗 Huggingface Deaftest Dataset] [🏆 Leaderboard]
🔥 News
2024.11.24 🌟 We release AV-Odyssey, the first-ever comprehensive evaluation benchmark to explore whether MLLMs really understand… See the full description on the dataset page: https://huggingface.co/datasets/AV-Odyssey/Deaftest_dataset.AV-QuantBench-Dataset
AV-QuantBench
AV-QuantBench is a procedural audio-visual benchmark for evaluating multimodal foundation models on abstract temporal reasoning, cross-modal conflict detection, and synchronized data interpretation across finance, medical, and industrial domains.
This Hugging Face dataset repository is structured as a benchmark-style release. It contains:
split metadata in JSONL format,
question-answer annotations,
audio-visual sample assets,
manifest files by domain,
and… See the full description on the dataset page: https://huggingface.co/datasets/gfcfirefly/AV-QuantBench-Dataset.mascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/electron-rare/mascarade-dsp-dataset.mascarade-dsp-dataset
Mascarade — DSP & Signal Processing Q&A
✅ ATTRIBUTION AUDIT COMPLETED (2026-05-11)
Per-sample Stack Exchange Electronics attribution recovered via the SE
/search/advanced + /questions/{id} API search :
169 samples (~5.35 %) confirmed as Stack Exchange Electronics
(CC-BY-SA-4.0) — fully attributed in metadata.stack_exchange_attribution
(URL + author display name + author user_id + post_id + creation_date_unix + match_confidence ≥ 0.60).
535 samples (~16.93 %) marked… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-dsp-dataset.
