datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShareGPT4Video
ShareGPT4Video 4.8M Dataset Card
Dataset details
Dataset type:
ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.
sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.SVCBench
SVCBench: Streaming Video Counting Benchmark
This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline.
Project Page: https://buaa-colalab.github.io/SVCBench/
Code: https://github.com/buaa-colalab/SVCBench
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/buaaplay/SVCBench.SITE-BenchThis dataset contains image and video QA test sets for SITE-Bench evaluation.
VideoHallucer
VideoHallucer
Paper: https://huggingface.co/papers/2406.16338
This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/bigai-nlco/VideoHallucer.SR-3D-Bench
Spatial Region 3D (SR-3D) Aware Benchmark
Paper: https://arxiv.org/abs/2509.13317Project page: https://www.anjiecheng.me/sr3dCode: https://github.com/AnjieCheng/SR-3D
[!IMPORTANT]
[Feb. 18, 2026] UPDATE: To improve compatibility with general-purpose VLMs, the benchmark is reformulated into multiple-choice and numerical questions following the VSI-Bench evaluation protocol. Videos are annotated with set-of-marks to explicitly indicate regions. The benchmark will be compatible… See the full description on the dataset page: https://huggingface.co/datasets/a8cheng/SR-3D-Bench.SportsTime
SportsTime
SportsTime is a long-form sports video question answering benchmark for temporal compositional reasoning, accepted to ECCV 2026.
It contains 14,326 open-ended QA pairs with 50,000+ step-wise temporal evidence annotations across 1,575 videos and five team sports: basketball, American football, ice hockey, soccer, and volleyball.
Dataset
This Hugging Face dataset repository provides the annotation files, official train/test split, and video files.… See the full description on the dataset page: https://huggingface.co/datasets/Ustiniansy/SportsTime.youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.molmo2-moments
Molmo-2 Moments (M2M)
Long-video QA dataset where every question is anchored to a specific
[start, end] clip interval in seconds. Released alongside the ToolMerge
paper, "Decomposing Queries into Tool Calls for Long-Video Keyframe
Retrieval".
⚠️ Source videos & ownership
The videos/*.mp4 files in this repository were collected from YouTube.
We do not own these videos and claim no copyright over them. All rights to
the video content remain with the original… See the full description on the dataset page: https://huggingface.co/datasets/michalsr/molmo2-moments.agentvidbench
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline)
by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI)
Layout
.
├── README.md
├── questions.jsonl # 100 rows — one per question
├── videos.jsonl # 71 rows — one per unique video
├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.Seamless_Dummy_Dataset_Fixed_3
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
Ava-100
Empowering Agentic Video Analytics Systems with Video Language Models
[🖥️ Project Code] [📖 arXiv Paper] [📊 Dataset]
Introduction
AVA-100 is an ultra-long video benchmark specially designed to evaluate video analysis capabilities Avas-100 consists of 8 videos, each exceeding 10 hours in length, and includes a total of 120 manually annotated questions. The benchmark covers four typical video analytics scenarios: human daily activities, city walking, wildlife… See the full description on the dataset page: https://huggingface.co/datasets/iesc/Ava-100.VideoHallucer
VideoHallucer
Paper: https://huggingface.co/papers/2406.16338
This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/ov015/VideoHallucer.valor32k-avqa-v2
Valor32k-AVQA v2.0
Valor32k-AVQA v2.0 is an open-ended audio-visual question answering dataset and benchmark with 28,861 videos and 225,487 question-answer pairs in this Hugging Face release. Each question is annotated with a modality label (visual, audio, or audio-visual) and one of six categories: description, action, count, temporal, location, and relative-position.
Links
Paper: ACM Digital Library
Project page: inesriahi.github.io/valor32k-avqa-2
Code and… See the full description on the dataset page: https://huggingface.co/datasets/inesriahi/valor32k-avqa-v2.agentvidbench
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline)
by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI)
Layout
.
├── README.md
├── questions.jsonl # 100 rows — one per question
├── videos.jsonl # 71 rows — one per unique video
├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench.Inst-It-Dataset
Inst-IT Dataset: An Instruction Tuning Dataset with Multi-level Fine-Grained Annotations
introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning
🌐 Homepage | Code | 🤗 Paper | 📖 arXiv
Inst-IT Dataset Overview
We create a large-scale instruction tuning dataset, the Inst-it Dataset. To the best of our knowledge, this is the first dataset that provides fine-grained annotations centric on specific… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Dataset.OmniCoding
OmniCoding
A multimodal terminal-tool-use SFT/RL dataset. Each record is a
question + verifiable answer + media (video/audio/image) — the target
agent is expected to operate on the media via a Linux terminal (ffmpeg,
ffprobe, whisper, python, etc.) rather than a GUI.
Aggregated and filtered from four upstream sources, with a single unified
schema, global dedup, and category-balanced sampling.
Records
Source
n
Omnimodal-Agent-SFT-2K (RUC-NLPIR) — agentic… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/OmniCoding.ViMU
ViMU: Benchmarking Video Metaphorical Understanding
Qi Li, Xinchao Wang*
*Corresponding author
xML Lab, National University of Singapore
Our GitHub repository contains the evaluation scripts for ViMU, a benchmark for video metaphorical understanding. The code evaluates multimodal models on four tasks:
Open-ended interpretation (OE)
Evidence grounding (EG)
Rhetoric mechanism identification (RM)
Social value signal identification (SV)
Directory Structure
Expected… See the full description on the dataset page: https://huggingface.co/datasets/LIQIIIII/ViMU.VUDG
VUDG: A Dataset for Video Understanding Domain Generalization
VUDG is a benchmark dataset for evaluating domain generalization (DG) in video understanding. It contains 7,899 video clips and 36,388 high-quality QA pairs, covering 11 diverse visual domains, such as cartoon, egocentric, surveillance, rainy, snowy, etc. Each video is annotated with both multiple-choice and open-ended question-answer pairs, designed via a multi-expert progressive annotation pipeline using large… See the full description on the dataset page: https://huggingface.co/datasets/QLGalaxy/VUDG.WereBench
Anonymization
For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information.
WereBench
WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior with… See the full description on the dataset page: https://huggingface.co/datasets/Yuan4629/WereBench.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.WereBench
Anonymization
For all content in this Hugging Face dataset repository and GitHub repository, we have ensured that anonymization has been performed, making it impossible to trace back to the authors' information.
WereBench
WereBench is a benchmark dataset for evaluating language models in the Werewolf (similar to Mafia) social deduction setting. It focuses on human‑aligned strategic reasoning rather than only coarse metrics (e.g., win rate), aligning model behavior… See the full description on the dataset page: https://huggingface.co/datasets/n0nam4/WereBench.MineExplorer
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
Tianjie Ju · Yueqing Sun · Zheng Wu · Wei Zhang · Yaqi Huo · Xi Su · Qi Gu · Xunliang Cai · Gongshen Liu · Zhuosheng Zhang
Abstract
Multimodal large language models (MLLMs) have shown strong capabilities in perception,
reasoning, and action generation. However, their ability to sustain exploration in dynamic
open worlds remains unclear. Existing embodied and… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/MineExplorer.EMCompressEMCompress
A Benchmark for Endomorphic Multimodal Compression on Long Cooking Videos
📰 News
2026.05 🎉 Dataset + reproduction code released on HuggingFace & GitHub.
2026.04 📝 Paper accepted to ACL 2026 Findings.
🧠 About
EMCompress is the first benchmark dedicated to evaluating the Endomorphic Multimodal Compression (EMC) task: an endomorphic transformation F_EMC : (V, Q) → (v, q) that compresses a (video, question) pair into a shorter… See the full description on the dataset page: https://huggingface.co/datasets/LordUky/EMCompress.FineBadmintonBenchmark
FineBadmintonBenchmark
Fine-grained badminton video question answering benchmark.
Dataset Structure
hf_video_clips_qa/: one clip per QA item, named by video_uid (for example video_000001.mp4).
finebadmintonbenchmark/: annotation JSON files.
each item contains video_uid
each QA item maps to exactly one video clip through video_uid
Citation
@inproceedings{he2025finebadminton,
title={Finebadminton: A multi-level dataset for fine-grained badminton video… See the full description on the dataset page: https://huggingface.co/datasets/iLearn-Lab/FineBadmintonBenchmark.E2E_real_object
E2E Real Object Direction
A video-based benchmark for evaluating VideoLLMs' directional reasoning and object recognition on real-world objects.
Conditions
Condition
Question
Answer
Purpose
direction_only
"In which direction is the object moving?"
"Up"
Baseline direction recognition
direction_obj_in_q
"In which direction is the car moving?"
"Up"
Does naming the object help?
direction_obj_in_a
"In which direction is the object moving?"
"The car is moving up"… See the full description on the dataset page: https://huggingface.co/datasets/KHUjongseo/E2E_real_object.EgoLink2026
EgoLink2026
EgoLink 2026 evaluation benchmark for egocentric social reasoning.
Dataset Structure
Eval.jsonl: 4030 multiple-choice evaluation questions
data/: 1055 egocentric video files referenced by the path field in each sample
CP-Bench
Continuous Perception Benchmark (CP-Bench)
Overview
The Continuous Perception Benchmark (CP-Bench) is a diagnostic dataset designed to evaluate whether modern vision-language and multimodal models can integrate continuous visual information over time—an ability that is central to human visual perception but largely absent in contemporary architectures. Inspired by the continuous, stream-based nature of human vision, CP-Bench isolates the core requirement of maintaining… See the full description on the dataset page: https://huggingface.co/datasets/ContinuousPerceptionResearch/CP-Bench.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench-sample.BlackSwanSuite-MCQ
Black Swan Suite
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events Aditya Chinchure*, Sahithya Ravi*, Raymond Ng, Vered Shwartz, Boyang Li, Leonid Sigal (* equal) 🎉 Accepted at CVPR 2025 arXiv | Website
Dataset Information
Black Swan has three variants of questions. Please find the data in the appropriate repositories:
BlackSwanSuite-Gen -- link
BlackSwanSuite-MCQ (this)
BlackSwanSuite-YN -- link
This dataset contains questions for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/UBC-ViL/BlackSwanSuite-MCQ.E2E_20K_test
E2E 20K Synthetic Direction
A large-scale synthetic video benchmark for evaluating VideoLLMs' directional reasoning.20,000 videos of colored geometric shapes moving in four cardinal directions, with automatically generated MCQ annotations.
Directions
Direction
Description
up
Object moves toward the top of the frame
down
Object moves toward the bottom of the frame
left
Object moves toward the left of the frame
right
Object moves toward the right of the… See the full description on the dataset page: https://huggingface.co/datasets/KHUjongseo/E2E_20K_test.
