datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VideoChat3-LV116k
VideoChat3-LV116K
VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments.
The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.vivid-video-instructvideo-quality-scored
Image-to-Video Quality-Scored Clips
A collection of prompted image-to-video samples with quality-evaluation metadata.
Each sample pairs a first frame (the I2V conditioning image) with one or both
of:
a generated video produced by a video model from the first frame + prompt
an original clip (the reference/source video the prompt was authored around)
A subset of the samples also carry per-clip quality scores: an overall
quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.VideoKR-Train
VideoKR-Train
📄 ArXiv
| 💻 Code
| 🤗 Collection
About
This repository contains the VideoKR training data presented in VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding (ICML 2026 Spotlight).
VideoKR is the first large-scale training corpus specifically designed for knowledge- and reasoning-intensive video understanding. It contains 315K video reasoning examples over 145K newly collected, CC-licensed… See the full description on the dataset page: https://huggingface.co/datasets/minuzero/VideoKR-Train.AVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.Temporal-Logic-Video-Dataset
Temporal Logic Video (TLV) Dataset
Temporal Logic Video (TLV) Dataset
Synthetic and real video dataset with temporal logic annotation
Explore the GitHub »
NSVS-TL Project Webpage
·
NSVS-TL Source Code
Overview
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components:
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.VideoChat3-Academic2M
VideoChat3-Academic2M
VideoChat3-Academic2M is the academic video instruction data used by VideoChat3. It re-annotates public academic video datasets for video captioning, video question answering, and fine-grained motion understanding.
The dataset follows an evidence-grounded annotation enhancement pipeline. Short answers, option-only labels, and concise captions are rewritten into richer instruction-following responses that mention visible objects, actions, scenes, temporal… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-Academic2M.VBVR-Bench-Data
VBVR: A Very Big Video Reasoning Suite
Overview
Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture,
enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality.
Systematically studying video reasoning and its scaling behavior suffers from a lack of… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Bench-Data.VideoHallucer
VideoHallucer
Paper: https://huggingface.co/papers/2406.16338
This work introduces VideoHallucer, the first comprehensive benchmark for hallucination detection in large video-language models (LVLMs). VideoHallucer categorizes hallucinations into two main types: intrinsic and extrinsic, offering further subcategories for detailed analysis, including object-relation, temporal, semantic detail, extrinsic factual, and extrinsic non-factual hallucinations. We adopt an adversarial binary… See the full description on the dataset page: https://huggingface.co/datasets/bigai-nlco/VideoHallucer.video2mentalego4d-videoEgoCOT is a large-scale embodied planning dataset, which selected egocentric videos from the Ego4D dataset and corresponding high-quality step-by-step language instructions, which are machine generated, then semantics-based filtered, and finally human-verified.
For mored details, please visit EgoCOT_Dataset.
If you find this dataset useful, please consider citing the paper,
@article{mu2024embodiedgpt,
title={Embodiedgpt: Vision-language pre-training via embodied chain of thought}… See the full description on the dataset page: https://huggingface.co/datasets/wofmanaf/ego4d-video.Video-IFBench
Video-IFBench
This release contains the evaluation split used for the Video-IFBench main experiments.
Paper: Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Project page: https://alexios-hub.github.io/Video-IFBench/
Code: https://github.com/Alexios-hub/Video-IFBench
Video-MMLU
Video-MMLU Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: Video-MMLU Benchmark
Features
Benchmark Collection and Processing
Video-MMLU specifically targets videos that focus on theorem demonstrations and probleming-solving, covering mathematics, physics, and chemistry. The videos deliver dense information through numbers and formulas, pose significant challenges for video LMMs in dynamic OCR… See the full description on the dataset page: https://huggingface.co/datasets/Enxin/Video-MMLU.Video_Reality_Test
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:
real_hard: 100 samples.… See the full description on the dataset page: https://huggingface.co/datasets/kolerk/Video_Reality_Test.Video-Detailed-Caption
Video Detailed Caption Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: AuroraCap Model
Huggingface: VDC Benchmark
Huggingface: Trainset
Features
Benchmark Collection and Processing
We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.VideoGPT-plus_Training_Datasettoc_bench
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
TOC-Bench is a diagnostic benchmark for evaluating whether Video Large Language Models maintain object identity, state, persistence, and temporal relations throughout a video. It focuses on object-centric phenomena including occlusion, disappearance, reappearance, repeated events, event order, temporal location, duration, conditional state, and relative movement.
Anonymous-review notice. This… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-video-benchmark/toc_bench.Long-video-test-dataFiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/LIMinghan/FiVE-Fine-Grained-Video-Editing-Benchmark.video-SALMONN_2_testset
video-SALMONN 2 Benchmark
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.pySocial-IQ-Video
Copy of Social-IQ 2.0 Challenge
We are hiring collaborators to organize a similar challenge like Social-IQ 2.0. If you are interested in it, please contact us via xucao@pediamed.ai.
2d_dungeon_flier_video_balanced
2D Dungeon Flier Video: Balanced Causal Splits
This dataset is a split-safe, balanced augmentation of osazuwa/2d_dungeon_flier_video. It reuses all 10,000 source episodes exactly once and adds 3,100 episodes from the same simulator. There is no clip overlap across splits.
Each episode is a 14-second MP4 with 140 frames at 10 FPS and a stored resolution of 900 x 540 pixels. Matching NPZ files contain the nine-variable causal trace, action tokens, and intervention encoding.
Every… See the full description on the dataset page: https://huggingface.co/datasets/osazuwa/2d_dungeon_flier_video_balanced.Video_Reality_Test
Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/ziweix/Video_Reality_Test.Video_Reality_Test
Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/Dii2/Video_Reality_Test.charades-stage1_data
VideoSearch-R1 Charades-STA
This repository contains the prepared Charades-STA artifacts used by VideoSearch-R1.
Paper: VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
Project Page: https://mlvlab.github.io/VideoSearch-R1/
Repository: https://github.com/mlvlab/VideoSearch-R1
Stage 1 Cold Start SFT
Charades_STA_Stage1_ColdStart.jsonl is the Stage 1 cold-start SFT dataset used to train the VideoSearch-R1 verifier/reasoner on… See the full description on the dataset page: https://huggingface.co/datasets/VideoSearchR1/charades-stage1_data.VideoKR-Eval
VideoKR-Eval
📄 ArXiv
| 💻 Code
| 🤗 Collection
About
This repository contains the VideoKR-Eval benchmark presented in VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding (ICML 2026 Spotlight).
VideoKR-Eval is an expert-annotated evaluation benchmark for knowledge- and reasoning-intensive video understanding. Unlike existing benchmarks where a substantial fraction of questions can be answered from a single… See the full description on the dataset page: https://huggingface.co/datasets/minuzero/VideoKR-Eval.VideoArgusBench
VideoArgusBench
Sample-specific rubric benchmark for conditioned video generation.
VideoArgusBench is the evaluation benchmark for VideoArgus, a framework that scores a generated
video against a rubric written for that specific prompt rather than a fixed global metric. This
dataset ships the inputs (conditioning assets + prompts) and, for each input, a rubric. It does
not contain generated videos — you bring your own model's outputs and score them with the VideoArgus
evaluation… See the full description on the dataset page: https://huggingface.co/datasets/zengziyun/VideoArgusBench.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.VideoInstruct-100KVideoInstruct100K is a high-quality video conversation dataset generated using human-assisted and semi-automatic annotation techniques. The question answers in the dataset are related to,
Video Summariazation
Description-based question-answers (exploring spatial, temporal, relationships, and reasoning concepts)
Creative/generative question-answers
For mored details, please visit Oryx/VideoChatGPT/video-instruction-data-generation.
If you find this dataset useful, please consider citing the… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/VideoInstruct-100K.
