datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
h3-fewstep-benchmark-20260920
H3 few-step benchmark
Fifteen few-step checkpoints for the MiniMax-H3 video+audio model, each run on the
same 48 English prompts × 2 seeds (0 and 1) × 3 step counts (4, 8, 32) =
288 videos per checkpoint, 4,320 videos in total. Conditioning is text only
(no input image, no camera trajectory). Every video is 1344×768, 5 seconds
(120 frames at 24 fps), with synthesized audio unless noted.
All checkpoints are adapters or distilled variants of MiniMaxAI/MiniMax-H3, except the… See the full description on the dataset page: https://huggingface.co/datasets/hffordata/h3-fewstep-benchmark-20260920.Video-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.GOAI-2026MOVA_benchmark_for_arena
MOVA Benchmark for Arena
This is the benchmark used for the subjective arena experiments of MOVA (MOVA: Towards Scalable and Synchronized Video–Audio Generation). All prompts are rewritten by the workflow introduced in the paper.
Paper: MOVA: Towards Scalable and Synchronized Video–Audio Generation
Code: https://github.com/OpenMOVA/MOVA
Overview
The benchmark contains 732 samples in total, organized into two subsets:
Subset
Samples
MOVA-Bench
132… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuzhang-0212/MOVA_benchmark_for_arena.benchmark-datasets
Latency-Sensitive Bench datasets
Accepted zero-latency teacher rollouts for the supported benchmark tasks.
Viewer subsets
humanoidbench_balance_simple: 90 training and 10 validation episodes. The
observation.image values are PNG bytes declared as the Hugging Face Image
feature, so the Dataset Viewer renders them instead of showing their encoded
representation. Canonical LeRobot MP4 files remain under each split's
videos/ directory.
mikasa_intercept_grab_fast:… See the full description on the dataset page: https://huggingface.co/datasets/latency-sensitive-bench/benchmark-datasets.MUMA-TOM-BENCHMARK
MuMA-ToM: Multi-modal Multi-Agent Theory of Mind AAAI 2025 (Oral)
[🏠Homepage] [💻Code] [📝Paper]
MuMA-ToM is the first multi-modal Theory of Mind benchmark designed to evaluate mental reasoning in embodied multi-agent interactions. The benchmark was designed with several key features in mind:
It is factually correct, concise, and readable.
It requires integrating information from multiple modalities to answer the questions.
It tests understanding of multi-agent interactions… See the full description on the dataset page: https://huggingface.co/datasets/SCAI-JHU/MUMA-TOM-BENCHMARK.qualcomm-exercise-video-dataset-benchmark
Dataset Card for Qualcomm Exercise Video Dataset (Benchmark)
This is the benchmark split of the dataset as described here
This is a FiftyOne dataset with 74 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/qualcomm-exercise-video-dataset-benchmark.VACE-Benchmark
VACE: All-in-One Video Creation and Editing
(ICCV 2025)
Zeyinzi Jiang*
·
Zhen Han*
·
Chaojie Mao*†
·
Jingfeng Zhang
·
Yulin Pan
·
Yu Liu
Tongyi Lab -
Introduction
VACE is an all-in-one model designed for video creation and editing. It encompasses various tasks, including reference-to-video generation (R2V), video-to-video editing (V2V), and masked video-to-video editing… See the full description on the dataset page: https://huggingface.co/datasets/ali-vilab/VACE-Benchmark.scene-mem-benchmark
scene-mem-benchmark
A benchmark for scene memory in embodied agents: an agent watches a mobile manipulator work
in a house for several minutes, then is asked to retrieve an object it has to remember — one
that was moved, dropped, or merely seen along the way — or (resume) to go back and finish the
job it was interrupted in, remembering how far it had got — or (routine) to put a new object away
where this household keeps that kind of thing, a rule it was never told and can only… See the full description on the dataset page: https://huggingface.co/datasets/Keh0t0/scene-mem-benchmark.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/rtx-5090-benchmarks.genvsr-video-benchmarksreacthuman-benchmark-scaled-with-videos
ReactHuman Benchmark — Scaled
A physics-grounded benchmark of household hazard scenarios for evaluating
embodied reactive decision-making. Each scene renders an object undergoing a
physical event (falling, tipping, thrown, bouncing, …) toward an observer; the
ground-truth action label (EXECUTE_CATCH / TRIGGER_DODGE /
BRACE_FOR_IMPACT) is derived from object properties, not speed.
Generated in LLM mode driving a procedural physics randomizer: Claude routes
each natural-language… See the full description on the dataset page: https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled-with-videos.VSI-Super-Wild-Benchmark
VSI-Super-Wild
This repository stores the public video assets and the current lmms-eval benchmark release for Toward Supersensing.
Current benchmark release
Version: cambw_v2_recheck_20260409_contentfix_mc4
Location: benchmarks/cambw_v2_recheck_20260409_contentfix_mc4/
QA count: 12021
Split count:
part1_long: 511
part2_3_short: 11510
Layout
videos/: public video files
benchmarks/cambw_v2_recheck_20260409_contentfix_mc4/data/: materialized… See the full description on the dataset page: https://huggingface.co/datasets/Gradygu3u/VSI-Super-Wild-Benchmark.scene-extrapolation-benchmark
Scene-Extrapolation Benchmark Suite
A unified, honest evaluation of generative novel-view synthesis (NVS) in the regimes where
per-scene 3D-Gaussian-Splatting reconstruction fails:
Extrapolative — held-out, unobserved views (not interpolation between dense captures).
Dynamic — moving foreground content.
Long-horizon — chained / loop-closing camera trajectories where drift accumulates.
Memory / revisit — recall when the camera returns to a previously-observed pose (do methods… See the full description on the dataset page: https://huggingface.co/datasets/luuuulinnnn/scene-extrapolation-benchmark.svi-benchmark
Stable Video Infinity (SVI) Benchmark Dataset
This benchmark dataset is introduced in the paper:
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
by Wuyang Li, Wentao Pan, Po-Chien Luan, Yang Gao, Alexandre Alahi (2025).
Project page: https://stable-video-infinity.github.io/homepage/
Code: https://github.com/vita-epfl/Stable-Video-Infinity
Abstract
We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with… See the full description on the dataset page: https://huggingface.co/datasets/epfl-vita/svi-benchmark.toc_bench
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models
TOC-Bench is a diagnostic benchmark for evaluating whether Video Large Language Models maintain object identity, state, persistence, and temporal relations throughout a video. It focuses on object-centric phenomena including occlusion, disappearance, reappearance, repeated events, event order, temporal location, duration, conditional state, and relative movement.
Anonymous-review notice. This… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-video-benchmark/toc_bench.reacthuman-benchmark-scaled
ReactHuman Benchmark — Scaled
A physics-grounded benchmark of household hazard scenarios for evaluating
embodied reactive decision-making. Each scene renders an object undergoing a
physical event (falling, tipping, thrown, bouncing, …) toward an observer; the
ground-truth action label (EXECUTE_CATCH / TRIGGER_DODGE /
BRACE_FOR_IMPACT) is derived from object properties, not speed.
Generated in LLM mode driving a procedural physics randomizer: Claude routes
each natural-language… See the full description on the dataset page: https://huggingface.co/datasets/Alan123/reacthuman-benchmark-scaled.robodyna-benchmark-v2
RoboDyna Benchmark
Expert demonstrations for RoboDyna, a dual-arm manipulation benchmark built around dynamic
scenes — moving targets, rolling and falling objects, closing time windows, conveyor belts, and
distractors — rather than static pick-and-place. Every episode is a scripted-expert rollout that
succeeded; failures are not published.
Built on RoboTwin 2.0 / DOMINO with SAPIEN 3.0.3 and a
dual-UR5 + WSG gripper embodiment (ur5-wsg).
At a glance… See the full description on the dataset page: https://huggingface.co/datasets/RoboDyna/robodyna-benchmark-v2.xplanner-benchmark
XPlanner-Benchmark
XPlanner-Benchmark is the portable release of the X-Planner 1,500-episode evaluation benchmark. It contains synchronized multi-view robot-manipulation videos and the episode-level task, subtask, action, scene, duration, and complexity metadata used by X-Planner.
Contents
1,500 episodes
3,490 MP4 video references
167 source dataset identifiers
525 unique task names
31 task classes
41 inferred atomic action labels
22.70 total hours of episode… See the full description on the dataset page: https://huggingface.co/datasets/x-square-robot/xplanner-benchmark.genvsr-video-benchmarks
GenVSR Video Benchmarks
Public video inputs and restoration outputs used by the GenVSR Visual Comparator.
Each immutable experiment version contains aligned MP4 sources, posters, and a manifest.
Please consult the individual version manifest for source labels and comparison defaults.
GeometryCrafter-benchmark
GeometryCrafter Benchmark Datasets
Preprocessed GeometryCrafter evaluation benchmark clips used for MotionCrafter evaluation.
Each sample typically contains:
*_rgb_320_640.mp4 — RGB video (320×640)
*_normed_data_320_640.hdf5 — normalized geometry / point-map data
meta_infos.txt — clip catalog (rgb_path data_path num_frames)
Subsets
Folder
Source
Approx. size
Kubric_video
Kubric
~3 GB
Spring_video
Spring
~14 GB
Virtual_KITTI_2_video
Virtual KITTI 2… See the full description on the dataset page: https://huggingface.co/datasets/MotionCrafterevaluation/GeometryCrafter-benchmark.assembly_benchmarkThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 821,
"total_frames": 620554,
"total_tasks": 75,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:821"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/pyrrosk/assembly_benchmark.WoW-1-Benchmark-Samples
🧠 WoW-1 Benchmark Samples
WoW-1 Benchmark Samples is the official evaluation dataset released as part of the WoW (World-Omniscient World Model) project. This benchmark is designed to assess the physical consistency and causal reasoning capabilities of generative world models for robotics and embodied AI.
📘 Dataset Overview
This dataset contains 612 natural language prompts representing real-world robot interaction tasks. These instructions are used to evaluate world… See the full description on the dataset page: https://huggingface.co/datasets/X-Humanoid/WoW-1-Benchmark-Samples.ebim_task2_realrobotdatavideo-full-duplex-benchmark
VideoFDB: Video-Full-Duplex-Benchmark
Project Page · HuggingFace · Paper (arXiv)
Dataset Description
A benchmark dataset of annotated, two-person video conference recordings designed to support the evaluation of multimodal AI agents in conversational settings. The dataset covers 11 distinct conversational dynamics — spanning verbal, nonverbal, and mixed-modality behavior — annotated through a three-pass human-in-the-loop pipeline.
The benchmark consists of trimmed… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/video-full-duplex-benchmark.robomind_benchmark1_1_release_franka_3rgb_mobile_marco_cup
benchmark1_1_release_franka_3rgb_mobile_marco_cup
This dataset converts the Robomain format uniformly into LeRobot V3.0.
Dataset Statistics
本体: franka_3rgb
末端执行器: 夹爪
任务平台显示版: 移动马克杯
total_episodes: 250
total_tasks: 1
size: 483.9M
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── images
│ └── observation.images.camera_top
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/robomind_benchmark1_1_release_franka_3rgb_mobile_marco_cup.LikePhys-BenchmarkAIGC_LipSync_Benchmark
AIGC-LipSync Benchmark
📋 Overview
AIGC-LipSync Benchmark is a comprehensive evaluation benchmark specifically designed for lip synchronization in AI-Generated Content (AIGC). This benchmark consists of 615 high-quality videos covering a wide spectrum of visual representations, from realistic humans to stylized characters, enabling thorough assessment of lip synchronization methods across diverse AI-generated video scenarios.
🎯 Key Features
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ZiqiaoPeng/AIGC_LipSync_Benchmark.Acmmm2025_video_benchmarkgameplay-benchmark-strict
Strict Gameplay Benchmark
This public release contains 1079 five-second gameplay clips
from 303 canonical game rows. Every clip is 1280x720, 30 FPS,
150 frames, H.264, and passed the strict black-frame, geometry, and dense content
gates. Videos are stored without recompression in benchmark_videos.zip.
benchmark_annotations.json contains one benchmark_clip annotation per archived
video and maps every record to its published_video_path. Source-specific license
review metadata is… See the full description on the dataset page: https://huggingface.co/datasets/shuaishuaicdp/gameplay-benchmark-strict.
