datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamingBench
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/mjuicem/StreamingBench.AV-SpeakerBench
AV-SpeakerBench
Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning.
Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/
Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench
Paper: https://arxiv.org/abs/2512.02231
Files
test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.SVBench
Dataset Card for SVBench
This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository.
Dataset Details
Dataset Description
SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of existing video… See the full description on the dataset page: https://huggingface.co/datasets/yzy666/SVBench.CogStream
CogStream Dataset
Dataset for CogStream: Context-guided Streaming Video Question Answering.
Overview
CogStream is a streaming video QA dataset designed to evaluate context-guided video reasoning. Models must identify and utilize relevant historical context to answer questions about ongoing video streams.
Statistics:
Split
Videos
QA Pairs
Train
852
55,623
Test
236
15,364
Total
1,088
70,987
Sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%)… See the full description on the dataset page: https://huggingface.co/datasets/SII-KYW/CogStream.GameplayQA
GameplayQA: A Decision-Dense POV-Synced Multi-Video
Understanding Benchmark of 3D Virtual Agents
Yunzhe Wang
Runhui Xu
Kexin Zheng
Tianyi Zhang
Jayavibhav N. Kogundi
Soham Hans
Volkan Ustun
University of Southern California
ACL 2026
Corresponding Author: yunzhewa@usc.edu
Overview
GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.MarineLife-16K
Dataset Card for MarineLife-16K
We introduce MarineLife-16K, a marine-domain video benchmark designed to evaluate the video understanding capabilities of Vision-Language Models (VLMs). MarineLife-16K contains 2,000 video-text pairs and 16,080 video-question-answer pairs across a collection of 2,000 marine videos, including 12,080 multiple-choice questions and 4,000 open-ended questions. The benchmark emphasizes specialized marine knowledge, visual reasoning, temporal… See the full description on the dataset page: https://huggingface.co/datasets/MarineLife-16K/MarineLife-16K.EgoAVU_data
[CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding
See our github for the code and setup instructions.
Check out our homepage, paper (CVPR) and paper (ICASSP) for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.StreamingBench-Slice
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
🏠 Project Page |
📄 arXiv Paper |
📦 Dataset |
🏅Leaderboard
StreamingBench evaluates Multimodal Large Language Models (MLLMs) in real-time, streaming video understanding tasks. 🌟
[NEW! 2025.05.15] 🔥: Seed1.5-VL achieved ALL model SOTA with a score of 82.80 on the Proactive Output.
[NEW! 2025.03.17] ⭐: ViSpeeker achieved Open-Source SOTA with a score of 61.60 on the… See the full description on the dataset page: https://huggingface.co/datasets/jayzhu486/StreamingBench-Slice.AudioVisual-Benchmark-Evaluation
AudioVisual Benchmark Evaluation — evaluation subsets
Item-id lists for the audio-visual benchmark subsets used in our reported
evaluation tables.
Layout
<benchmark>/eval_subset.csv item ids evaluated in the paper
<benchmark>/media_index.csv id -> media filename(s)
<benchmark>/media/ the media files those ids refer to
eval_subset.csv holds a single id column keyed to the source benchmark
(question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.jurisbenchomni
JurisBenchOmni — Model Predictions
This repository hosts the per-sample model predictions and the
analysis-ready aggregated CSVs that accompany our paper JurisBenchOmni:
A Judicial-Practice-Oriented Benchmark for Legal Omni-Modal Evaluation.
What you get here lets you read off, slice, and re-aggregate the
per-task / per-dimension / pipeline-level numbers we cite in the paper,
without having to re-run inference. Companion code that produces the
aggregated tables from the per-sample… See the full description on the dataset page: https://huggingface.co/datasets/jurisbenchomni-anonymous/jurisbenchomni.SVBench
Dataset Card for SVBench
This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources. For details, see our Project, Paper and GitHub repository.
Dataset Details
Dataset Description
SVBench is the first benchmark specifically designed to evaluate long-context streaming video understanding through temporal multi-turn question-answering (QA) chains. It addresses the limitations of… See the full description on the dataset page: https://huggingface.co/datasets/wulongcham/SVBench.morse-500
MORSE-500 Benchmark
🔥 News
May 15, 2025: We release MORSE-500, 500 programmatically generated videos across six reasoning categories: abstract, mathematical, physical, planning, spatial, and temporal, to stress-test multimodal reasoning. Frontier models including OpenAI o3 and Gemini 2.5 Pro score lower than… See the full description on the dataset page: https://huggingface.co/datasets/video-reasoning/morse-500.EgoAVU_data
[CVPR2026] EgoAVU
Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding
See our github for the code and setup instructions.
Check out our homepage and paper for more information.
We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual understanding. EgoAVU enriches existing egocentric narrations by integrating human actions with environmental context, explicitly linking visible objects and the sounds produced during interactions… See the full description on the dataset page: https://huggingface.co/datasets/jun111111/EgoAVU_data.
