datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ABot-World-Explorer-500h
ABot World Explorer 500h
ABot World Explorer 500h contains 30,969 action-conditioned video episodes
associated with the data infrastructure described in
ABot-World-0. Each episode preserves an MP4,
dataset-native keyboard actions, captions, and one COLMAP text sparse model.
Dataset facts
Item
Value
Episodes
30,969
Source objects
185,814
Semantic splits
None
License
Apache-2.0
The repository name is an identifier, not an audited… See the full description on the dataset page: https://huggingface.co/datasets/acvlab/ABot-World-Explorer-500h.robovqaVideoChat3-LV116k
VideoChat3-LV116K
VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video data with supervision over longer temporal contexts, where evidence can be sparse, delayed, and distributed across multiple video segments.
The dataset is constructed through a long-video synthesis pipeline. Candidate long videos are filtered for visual quality, semantic content, and temporal coherence. Videos are then split into manageable… See the full description on the dataset page: https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k.Vript
🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo]
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips). The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning, tilting, etc).… See the full description on the dataset page: https://huggingface.co/datasets/Mutonix/Vript.M3_VOS
[CVPR 2025] M3-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Venue: CVPR2025
Repository: 🛠️Tool, 🏠Page
Paper: arxiv.org/html/2412.13803v2
Point of Contact: Jiaxin Li , Zixuan Chen
📁 Structure
This dataset contains annotated videos and images for object segmentation tasks with phase transition information. The directory… See the full description on the dataset page: https://huggingface.co/datasets/Lijiaxin0111/M3_VOS.CG-Bench
CG-Bench
Project Website: https://cg-bench.github.io/leaderboard/GitHub Repository: https://github.com/CG-Bench/CG-Bench (includes running code)
Summary
We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question answering in long videos, addressing the limitations of existing benchmarks that focus primarily on short videos and rely on multiple-choice questions (MCQs). These limitations allow models to answer by elimination rather than genuine… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-Bench.InternData-fractal20220817_dataShareGPT4Video
ShareGPT4Video 4.8M Dataset Card
Dataset details
Dataset type:
ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos.
It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora.
sharegpt4video_40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page: https://huggingface.co/datasets/ShareGPT4Video/ShareGPT4Video.RoboInter-Data
RoboInter-Data: Intermediate Representation Annotations for Robot Manipulation
Rich, dense, per-frame intermediate representation annotations for robot manipulation, built on top of DROID and RH20T. Developed as part of the RoboInter project. You can try our Online demo.
The annotations cover 230k episodes and include: subtasks,
primitive skills, segmentation, gripper/object bounding boxes, placement proposals, affordance boxes,
grasp poses, traces, contact points, etc. And each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/RoboInter-Data.VOST-TAS
[NeurIPS 2025] Tracking and Understanding Object Transformations
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Dataset Visualizations: GitHub
Paper: arXiv:2511.04678
Project Page: tubelet-graph.github.io
Project Repository: GitHub
Point of Contact: Yihong Sun
📊 Dataset Overview
VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.minWM-dataMSR-VTTClone from "friedrichor/MSR-VTT".
MSRVTT contains 10K video clips and 200K captions.
We adopt the standard 1K-A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text-Video Retrieval field.
Train:
train_7k: 7,010 videos, 140,200 captions
train_9k: 9,000 videos, 180,000 captions
Test:
test_1k: 1,000 videos, 1,000 captions
🌟 Citation
@inproceedings{xu2016msrvtt,
title={Msr-vtt: A large video description dataset… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSR-VTT.vivid-video-instructAIGVDBenchvideo-quality-scored
Image-to-Video Quality-Scored Clips
A collection of prompted image-to-video samples with quality-evaluation metadata.
Each sample pairs a first frame (the I2V conditioning image) with one or both
of:
a generated video produced by a video model from the first frame + prompt
an original clip (the reference/source video the prompt was authored around)
A subset of the samples also carry per-clip quality scores: an overall
quality_score, six per-aspect breakdowns… See the full description on the dataset page: https://huggingface.co/datasets/mohantesting/video-quality-scored.SparseVideoNav
SparseVideoNav Datasets
This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav:
BVN: Beyond-the-View Navigation.
IFN: Instruction-Following Navigation.
Project links:
Project page: https://opendrivelab.com/SparseVideoNav
GitHub: https://github.com/OpenDriveLab/SparseVideoNav
Paper: https://arxiv.org/abs/2602.05827
Dataset Summary
SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.ProLongVid_data
Dataset Card for ProLongVid-data
Uses
This dataset is used for the training of the ProLongVid model. We only allow the use of this dataset for academic research and education purpose.
Paper: For more details, please check our paper
Code: For training recipe and other update, please refer to github repo.
Citation
@inproceedings{
wang2025prolongvid,
title={ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning},
author={Rui Wang… See the full description on the dataset page: https://huggingface.co/datasets/prolongvid/ProLongVid_data.DiDeMoClone from friedrichor/DiDeMo.
About
DiDeMo contains 10K long-form videos from Flickr. For each video, ~4 short sentences are annotated in temporal order. We follow the existing works to concatenate those short sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 8,395 videos, 8,395 captions (concatenate from 33,005 short captions)
Val: 1,065 videos, 1,065 captions (concatenate from 4,290 short captions) (We don't have… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/DiDeMo.S-EMBER
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
Episodic-memory video QA benchmark (face-blurred, audio-removed).
License & usage
This dataset is licensed under
CC BY-NC 4.0 and is provided
for non-commercial research use only. Access is gated: you must accept the
non-commercial terms above before downloading.
Contents
sember_mcq.jsonl — multiple-choice evaluation split.
sember_grounding.jsonl — answer-generation and… See the full description on the dataset page: https://huggingface.co/datasets/facebook/S-EMBER.MSVDClone from "friedrichor/MSVD".
MSVD contains 1,970 videos, each of which is paired with ~40 captions.
We adopt the official split:
Train: 1,200 videos, 48,774 captions
Val: 100 videos, 4,290 captions
Test: 670 videos, 27,763 captions
🌟 Citation
@inproceedings{chen2011collecting,
title={Collecting highly parallel data for paraphrase evaluation},
author={Chen, David and Dolan, William B},
booktitle={Proceedings of the Annual Meeting of the Association for… See the full description on the dataset page: https://huggingface.co/datasets/VLM2Vec/MSVD.OraRL-Data
OraRL-Data
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code]
We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL.
It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.MMVU
MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
🌐 Homepage •
🥇 Leaderboard •
📖 Paper •
🤗 Data
📰 News
2025-01-21: We are excited to release the MMVU paper, dataset, and evaluation code!
👋 Overview
Why MMVU Benchmark?
Despite the rapid progress of foundation models in both text-based and image-based expert reasoning, there is a clear gap in evaluating these models’ capabilities in specialized-domain video understanding.… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/MMVU.QAEgo4D-MC-testThis benchmark was collected by QAEgo4D and updated by GroundVQA.
We conducted some processing for the experiments presented in our paper ReKV.
SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.EgoExoLearnNOTE: Videos in huggingface are unprocessed, full-size videos. For benchmark and gaze alignment, we use processed 25fps videos. For processed data and code for benchmark, please visit the github page.
EgoExoLearn
This repository contains the video data of the following paper:
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World
Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong… See the full description on the dataset page: https://huggingface.co/datasets/hyf015/EgoExoLearn.Music-AVQASVCBench
SVCBench: Streaming Video Counting Benchmark
This dataset contains the clipped video segments for SVCBench, a Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance. It repositions counting as a minimal, controlled probe for diagnosing how video understanding models maintain world state along the video timeline.
Project Page: https://buaa-colalab.github.io/SVCBench/
Code: https://github.com/buaa-colalab/SVCBench
Dataset Description
This… See the full description on the dataset page: https://huggingface.co/datasets/buaaplay/SVCBench.AVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.
