datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Video-MMEshort_video_ocr_dataset
Short Video OCR / ASR Dataset
An actively curated research dataset for building OCR, ASR, subtitle-alignment,
and video-transcript pipelines for short social videos. It combines source
videos and extracted frames with human review artifacts and model-generated
text candidates. The primary languages are Ukrainian and Russian; English or
mixed-language content may also occur.
Status: work in progress. Model outputs and pseudo-label candidates are
not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.Video-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.VideoUFO
News
🎉 Accepted by NeurIPS 2025 Datasets and Benchmarks Track.
✨ Ranked Top 1 in the Hugging Face Dataset Trending List for text-to-video generation on March 7, 2025.
💗 Financially supported by OpenAI through the Researcher Access Program.
🌟 Downloaded 10,000+ times on Hugging Face after one month of release.
🔥 Featured in Hugging Face Daily Papers on March 4, 2025.
👍 Drawn interest from leading companies such as Alibaba Group and Baidu Inc.
😊 Recommended by a famous writer on… See the full description on the dataset page: https://huggingface.co/datasets/WenhaoWang/VideoUFO.vlabench_composite_ft_lerobot_videovlabench_primitive_ft_lerobot_video
VLABench Primitive Tasks Dataset - LeRobot v3.0
Dataset Description
This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially.
Compared with the v2.0 version and the RLDS version of the dataset, this release stores visual observations in a video-compressed format rather than as individual image files. This design provides significant advantages in both storage efficiency and data loading performance.… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_ft_lerobot_video.VideoUFO
News
✨ Ranked Top 1 in the Hugging Face Dataset Trending List for text-to-video generation on March 7, 2025.
Summary
This is the dataset proposed in our paper VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
VideoUFO is the first dataset curated in alignment with real-world users’ focused topics for text-to-video generation. Specifically, the dataset comprises over 1.09 million video clips spanning 1,291 topics. Here, we select the top 20 most… See the full description on the dataset page: https://huggingface.co/datasets/VideoUFO/VideoUFO.VideoGameQA-Bench
VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
by Mohammad Reza Taesiri, Abhijay Ghildyal, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer
Abstract:
With video games now generating the highest revenues in the entertainment industry, optimizing game development workflows has become essential for the sector's sustained growth. Recent advancements in Vision-Language Models (VLMs) offer considerable potential to automate and… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/VideoGameQA-Bench.xd-violence-rgb-videomae-chunked-testVideoDetailCaptionVideo-MMEVideoMMMUThis dataset contains the data for the paper Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos. Video-MMMU is a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos.
Project page: https://videommmu.github.io/
Leaderboard (last updated: 07 Feb, 2025)
Model
Overall
Perception
Comprehension
Adaptation
Δknowledge
Human Expert
74.44
84.33
78.67
60.33
+33.1… See the full description on the dataset page: https://huggingface.co/datasets/lmms-eval/VideoMMMU.VideoEval-Pro
VideoEval-Pro
VideoEval-Pro is a robust and realistic long video understanding benchmark containing open-ended, short-answer QA problems. The dataset is constructed by reformatting questions from four existing long video understanding MCQ benchmarks: Video-MME, MLVU, LVBench, and LongVideoBench into free-form questions. The paper can be found here.
The evaluation code and scripts are available at: TIGER-AI-Lab/VideoEval-Pro
Dataset Structure
Each example in the… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VideoEval-Pro.VideoMarathon
Dataset Card for VideoMarathon
VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately 9,700 hours, comprising 3.3 million QA pairs across 22 task categories.
Paper and more resources: [arXiv] [Project Website] [GitHub] [Model]
Intended Uses
This dataset is used for academic research purposes only.
Task Taxonomy
The dataset contains 22 diverse tasks over six fundamental topics, including temporality… See the full description on the dataset page: https://huggingface.co/datasets/GarrickAI/VideoMarathon.VBVR-Pro-SFT-Image
VBVR-Pro-SFT-Image
The interleaved-image supervised-fine-tuning split of VBVR-Pro: 1.24M programmatically generated reasoning instances across 250 parameterized tasks, one tar.gz per task.
Where VBVR-Pro-SFT-Video asks a model to render the reasoning process as a… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-SFT-Image.Video_logvideo-benchmark-resultsVBVR-Pro-SFT-Video
VBVR-Pro-SFT-Video
The video (I2V) supervised-fine-tuning split of VBVR-Pro: 1.24M programmatically generated reasoning instances across 250 parameterized tasks, one tar.gz per task.
At a glance
Property
Value
Tasks
250
Instances
1,250,000… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-SFT-Video.VBVR-Pro-RL
VBVR-Pro-RL
The reinforcement-learning split of VBVR-Pro: 50 parameterized tasks × 1,000 instances, held out from the SFT splits, in both a video (TI2V) and an interleaved-image setting.
At a glance
Property
Value
Tasks
50
Instances per… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Pro-RL.text-2-video-human-preferences
Rapidata Video Generation Preference Dataset
This dataset was collected in ~12 hours using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
The data collected in this dataset informs our text-2-video model benchmark. We just started so currently only two models are represented in this set:
Sora
Hunyouan
Pika 2.0
Runway ML Alpha
Luma Ray 2
Explore our latest model rankings on our website.
If you get value from this dataset and would… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences.text-2-video-human-preferences-wan2.1
Rapidata Video Generation Alibaba Wan2.1 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.Video2Reaction
Video2Reaction
Video2Reaction (V2R) is a multimodal dataset that maps short movie segments to the
distributional induced emotional reactions of viewers in the wild, as expressed through
social media comments. Unlike datasets that capture perceived emotion (the emotion
expressed by on-screen characters or filmmaker intent), Video2Reaction targets induced
emotion — the emotional response actually elicited in the audience — and represents each
clip's reaction as a probability… See the full description on the dataset page: https://huggingface.co/datasets/infofusionlab/Video2Reaction.VideoFeedback📃Paper | 🌐Website | 💻Github | 🛢️Datasets | 🤗Model | 🤗Demo
Overview
VideoFeedback contains a total of 37.6K text-to-video pairs from 11 popular video generative models,
with some real-world videos as data augmentation.
The videos are annotated by raters for five evaluation dimensions:
Visual Quality, Temporal Consistency, Dynamic Degree,
Text-to-Video Alignment and Factual Consistency, in 1-4 scoring scale.
VideoFeedback is used to for trainging of VideoScore
Below we… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VideoFeedback.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.TaskMeAnything-v1-videoqa-2024
Dataset Card for TaskMeAnything-v1-videoqa-2024
TaskMeAnything-v1-videoqa-2024 benchmark dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-2024-Videoqa
TaskMeAnything-v1-videoqa-2024 is a benchmark for reflecting the current progress of MLMs by automatically finding tasks that SOTA MLMs struggle with using the TaskMeAnything Top-K queries.
This… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-videoqa-2024.TaskMeAnything-v1-videoqa-random
Dataset Card for TaskMeAnything-v1-videoqa-random
TaskMeAnything-v1-videoqa-random dataset
🌐 Website | 📑 Paper | 🤗 Huggingface | 💻 Interface
If you like our project, please give us a star ⭐ on GitHub for latest update.
TaskMeAnything-v1-Random
TaskMeAnything-v1-videoqa-random is a dataset which randomly sampled questions from TaskMeAnything-v1, including 2,700 VideoQA questions. The dataset contains 9 splits, while each splits contains 300 questions… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/TaskMeAnything-v1-videoqa-random.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.VideoChatGPTvideo-tt
Towards Video Thinking Test (Video-TT): A Holistic Benchmark for Advanced Video Reasoning and Understanding
Video-TT comprises 1,000 YouTube videos, each paired with one open-ended question and four adversarial questions designed to probe visual and narrative complexity.
Paper: https://arxiv.org/abs/2507.15028
Project page: https://zhangyuanhan-ai.github.io/video-tt/
🚀 What's New
[2025.03] We release the benchmark!
1. Why Do We Need a New… See the full description on the dataset page: https://huggingface.co/datasets/lmms-eval/video-tt.Molmo2-VideoPoint
Molmo2-VideoPoint
Molmo2-VideoPoint is a dataset of video pointing data collected from human annotators.
It can be used to fine-tune vision-language models for video grounding by pointing.
Molmo2-VideoPoint is part of the Molmo2 dataset collection and was used to train the Molmo2 family of models.
Quick links:
📃 Paper
🎥 Blog with Videos
Usage
from datasets import load_dataset
# Load entire dataset
ds = load_dataset("allenai/Molmo2-VideoPoint", split="train")
#… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Molmo2-VideoPoint.
