datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIR-Bench-Dataset
AIR-Bench
Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks.
The former consists of 19 tasks with approximately 19k single-choice questions.
The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon).
Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.SciTS
SciTS: Scientific Time Series Understanding and Generation with LLMs
This repository contains the official dataset for SciTS: Scientific Time Series Understanding and Generation with LLMs (ICLR 2026). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances.
Dataset Structure
The benchmark is organized… See the full description on the dataset page: https://huggingface.co/datasets/ZJTustc/SciTS.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.SciTS
SciTS: Scientific Time Series Understanding and Generation with LLMs
This repository contains the official dataset for SciTS: Scientific Time Series Understanding and Generation with LLMs (ICLR 2026). SciTS is a large-scale benchmark designed to evaluate the capabilities of large language models on complex scientific time series data. It spans 12 scientific disciplines, 43 distinct tasks, and includes 54,023 instances.
Dataset Structure
The benchmark is organized… See the full description on the dataset page: https://huggingface.co/datasets/OpenTSLab/SciTS.Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.AudioMarathon
🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs
Abstract
AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars:
long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.human_behavior_atlas
Human Behavior Atlas
A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.
This dataset was used to train OmniSapiens, a foundation model for social behavior processing.
Papers:
Human Behavior Atlas: Benchmarking Unified Psychological and… See the full description on the dataset page: https://huggingface.co/datasets/HumanBehaviorAtlas/human_behavior_atlas.MMSU
[ICLR 2026] MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
Overview of MMSU
MMSU (Massive Multi-task Spoken Language Understanding and Reasoning Benchmark) is a comprehensive benchmark for evaluating fine-grained spoken language understanding and reasoning in multimodal models.
It systematically captures the variance of real-world linguistic phenomena in daily speech through 47 sub-tasks, including phonetics, prosody, rhetoric… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/MMSU.belebele-fleurs
Belebele-Fleurs
Belebele-Fleurs is a dataset suitable to evaluate two core tasks:
Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form.
Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.WearVox
WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
Paper: WearVox: An Egocentric Multichannel Voice Assistant Benchmark for WearablesAuthors: Zhaojiang Lin*, Yong Xu*, Kai Sun*, Jing Zheng, Yin Huang, Surya Appini, Krish Narang, Renjie Tao, Ishan Kapil Jain, Siddhant Arora, Ruizhi Li, Yiteng Huang, Kaushik Patnaik, Wenfang Xu, Suwon Shon, Yue Liu, Ahmed Aly, Anuj Kumar, Florian Metze, Luna DongAffiliations: Meta Reality Labs, Meta
📝 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zlinao/WearVox.OmniVideoBench
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
✨ Overview
Recent advances in multimodal large language models (MLLMs) have brought remarkable progress in video understanding.However, most existing benchmarks fail to jointly evaluate both audio and visual reasoning — often focusing on one modality or overlooking their interaction.
🎬 OmniVideoBench fills this gap.It’s a large-scale, rigorously curated… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/OmniVideoBench.VocalBench
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
This is the official release of VocalBench
Citation
If you find our work helpful, please cite our paper:
@article{liu2025vocalbench,
title={VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models},
author={Liu, Heyang and Wang, Yuhao and Cheng, Ziyang and Wu, Ronghua and Gu, Qunshan and Wang, Yanfeng and Wang, Yu},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/VocalNet/VocalBench.sib-fleurs
SIB-Fleurs
SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to.
The topics are:
Science/Technology
Travel
Politics
Sports
Health
Entertainment
Geography
Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*.
Dataset creation
This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.AudioJailbreak
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
AudioJailbreak is a benchmark framework specifically designed for evaluating the security of Audio Language Models (Audio LLMs). This project tests model defenses against malicious requests through various audio perturbation techniques.Note: This project aims to improve the security of audio language models. Researchers should use this tool responsibly.
📋 Table of Contents… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/AudioJailbreak.2M-Belebele
2M-Belebele
Highly-Multilingual Speech and American Sign Language Comprehension Dataset
We introduce 2M-Belebele as the first highly multilingual speech and American Sign Language (ASL) comprehension dataset. Our dataset, which is an extension of the existing Belebele only-text dataset, covers 74 spoken languages at the intersection of Belebele and Fleurs, and one sign language (ASL).
The speech dataset is built from aligning Belebele, Flores200 and Fleurs datasets as… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Belebele.AVQA
Summary | 摘要
This dataset is collected from the AVQA training subset (train_qa.json). We converted the data to the R1-AQA format, where each line in the text file represents a JSON object with specific keys.
The AVQA training set originally consists of approximately 40k samples. However, we use only about 38k samples because some data sources have become invalid (e.g. link failure, or less than 10 seconds).
Given that there is no quick link to the audio mentioned in the above two… See the full description on the dataset page: https://huggingface.co/datasets/Joysw909/AVQA.TimeGround-1M
TimeGround-1M
Synthetic English audio dataset for time-aware speech understanding, covering temporal localization, temporal description, and timed summaries.
Data Filtering
We use 14k hours of audio from YODAS2 English shards, selected from a 24k-hour source pool after language- and silence-ratio filtering. Synthetic annotations were generated for three time-grounded tasks, then filtered through LLM-based verification, deterministic validity checks, and… See the full description on the dataset page: https://huggingface.co/datasets/ai-sage/TimeGround-1M.Seamless_Dummy_Dataset_Fixed
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
MCIF
Dataset Description, Collection, and Source
MCIF (Multimodal Crosslingual Instruction Following) is a multilingual human-annotated benchmark
based on scientific talks that is designed to evaluate instruction-following in crosslingual,
multimodal settings over both short- and long-form inputs.
MCIF spans three core modalities -- speech, vision, and text -- and four diverse languages (English, German, Italian, and Chinese),
enabling a comprehensive evaluation of MLLMs'… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/MCIF.youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.OmniGAIA
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is designed to evaluate long-horizon, multi-hop, open-form problem solving in realistic settings rather than short perception-only QA.
Benchmark Construction
The OmniGAIA construction… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/OmniGAIA.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.LongAudioSpan
LongAudioSpan: Spanning the Duration and Depth of Audio Comprehension
Introduction
LongAudioSpan is a benchmark for long-form audio comprehension, spanning
diverse durations and cognitive depths.
Questions come from two complementary paths:
Native QA: questions drawn from the audio's natural content.
Anchor QA: questions built around acoustic anchors planted into the
audio.
Each path is scored in its own mode:
Accuracy: multiple choice… See the full description on the dataset page: https://huggingface.co/datasets/holvan/LongAudioSpan.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.dowis
Do What I Say (DOWIS): A Spoken Prompt Dataset for Instruction-Following
NEW DOWIS now also contains spoken and written prompts in Albanian (sq), and for the tasks LIPREAD and SLU!
TL;DR — DOWIS is a multilingual dataset of human-recorded spoken and written instruction prompts, designed to enable realistic evaluation of Speech Large Language Models across 11 tasks and 12 languages.
Dataset Summary
Most Speech LLM benchmarks use text-based prompts, which… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/dowis.human_behavior_atlas
Human Behavior Atlas
A large-scale multimodal dataset for human behavior understanding, spanning emotion recognition, sentiment analysis, humor detection, mental health screening, and video question answering. The dataset integrates 16 source datasets into a unified schema with audio, video, and pre-extracted features.
This dataset was used to train OmniSapiens, a foundation model for social behavior processing.
Papers:
Human Behavior Atlas: Benchmarking Unified Psychological… See the full description on the dataset page: https://huggingface.co/datasets/DennisDengHUst/human_behavior_atlas.sonic-o1
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
🎯 What is SONIC-O1?
The first open-source benchmark for evaluating omnimodal video understanding with systematic fairness analysis. SONIC-O1 requires models to jointly process audio, video, and social context from real-world interactions—not just transcripts.
Key… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/sonic-o1.Light-Omni-Training
Light-Omni Training Dataset
This repository contains the training data used by Light-Omni, a multimodal
agent framework for reflexive video understanding with long-term memory.
Light-Omni uses memory-augmented multimodal streams to train adapters for
memory construction, response generation, and reaction/action control.
Links
Project page: https://clare-nie.github.io/Light-Omni/
Code: https://github.com/Clare-Nie/Light-Omni
Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.MECAT-QAMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-Caption (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-QA.
