datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.OpenVid-1M
Summary
This is the dataset proposed in our paper [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.
OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets.
All videos in the OpenVid-1M dataset have resolutions of at least 512×512.… See the full description on the dataset page: https://huggingface.co/datasets/nkp37/OpenVid-1M.FreeTacMan
📦 FreeTacman
Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation [ICRA 2026]
🎯 Overview
This dataset supports the paper FreeTacman: Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation.
It contains a large-scale, high-precision visuo-tactile manipulation dataset with over 3000k visuo-tactile image pairs, more than 10k trajectories across 50 tasks.
We provide 🤗 Script (Hugging Face) and 👾 Script… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/FreeTacMan.10Kh-RealOmin-OpenDataBoasting over 10,000 hours of cumulative data and 1 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Compared with other datasets, it has the following advantages:
Ample Data Volume & Strong Generalization
Each skill is supported by sufficient data, collected from over 3,000 households and nearly 10,000 distinct fine-grained targets. It avoids simple repetitions and ensures robust generalization.
Authentic Scenarios & Focused… See the full description on the dataset page: https://huggingface.co/datasets/ad1t7a/10Kh-RealOmin-OpenData.Galaxea-Open-World-Dataset
Galaxea Open-World Dataset
Key Features
500+ hours of real-world mobile manipulation data.
All data collected using one uniform robotic embodiment (R1-Lite) for consistency.
Fine-grained subtask language annotations (bilingual Chinese/English).
Covers residential, kitchen, retail, and officesettings.
Dataset in LeRobot v2.1 format.
Dataset Structure
The dataset is organized as 227 task-level tar.gz archives under the lerobot/ directory. Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset.opendataLanguage: English (current) · 中文
Representative frames from TacVerse's bimanual
demonstrations.
Collected with XTac-UMI-G1 grippers, released as LeRobot
datasets.
TacVerse Open Data
Collection of 122 LeRobot v3.0 task datasets — 17,690 episodes,
370.2 hours, 40.0M frames, ~145 GB.
Each subfolder is a standalone LeRobot dataset (meta/info.json, data/, videos/).
Collection timestamps have been removed from titles and metadata.
Every frame carries six synchronized video… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/opendata.ExpVid
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
We present ExpVid, a benchmark to evaluate MLLMs on scientific experiment videos. ExpVid comprises 10 tasks across 3 levels, curated from a collection of 390 lab experiment videos spanning 13 disciplines.
How to Use
from datasets import load_dataset
dataset = load_dataset("OpenGVLab/ExpVid")
All task annotation .jsonl files are stored under annotations/level_*.
Each annotation includes the field:… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/ExpVid.SparseVideoNav
SparseVideoNav Datasets
This repository contains the real-world navigation datasets released with OpenDriveLab/SparseVideoNav:
BVN: Beyond-the-View Navigation.
IFN: Instruction-Following Navigation.
Project links:
Project page: https://opendrivelab.com/SparseVideoNav
GitHub: https://github.com/OpenDriveLab/SparseVideoNav
Paper: https://arxiv.org/abs/2602.05827
Dataset Summary
SparseVideoNav studies real-world vision-language navigation with sparse future… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/SparseVideoNav.OpenVid-1M-wds
OpenVid-1M — WebDataset repackaging
This repository is a sequential-read-optimized WebDataset repackaging of nkp37/OpenVid-1M by Nan et al. (ICLR 2025). The video content is identical to the original — only the on-disk layout is changed so it can be streamed efficiently from a single HTTP/NFS connection.
What differs from the original
Aspect
Original nkp37/OpenVid-1M
This repository
Format
Per-video mp4 files zipped
WebDataset .tar shards (~2 GB each)… See the full description on the dataset page: https://huggingface.co/datasets/Dev-Jahn/OpenVid-1M-wds.D2E-480p
D2E-480p
Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation
This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 268.7 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents.
What's included:
Video + Audio: H.264 encoded at 480p 60fps with game audio. Fixed 0.5s keyframe intervals and… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/D2E-480p.opencs2_dataset_wds
OpenCS2 - POV Renders WebDataset
Browse with the OpenCS2 Viewer - every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each
sample is one player's perspective for one round; ten POVs per round share the same tick clock.
Per POV round:
Video - 1280x720 @ 32 fps, near-lossless H.264, faststart, muxed with audio.
Audio - per-player stereo, mixed from that player's… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset_wds.assetsopencs2_dataset
OpenCS2 - POV Renders
Browse with the OpenCS2 Viewer - every match, map and round, with all 10 player POVs synced on one timeline.
Tick-aligned Counter-Strike 2 POV training clips, rendered from
blanchon/cs2_dataset_demo. Each row
in the main table is one player's perspective for one round; ten POVs per round share the same tick
clock.
Per POV round:
Video - 1280x720 @ 32 fps, near-lossless H.264, faststart, muxed with audio.
Audio - per-player stereo, mixed from that… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset.OpenVE-3M
OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
Haoyang He1*,
Jie Wang2*,
Jiangning Zhang1,
Zhucun Xue1,
Xingyuan Bu2,
Qiangpeng Yang2,
Shilei Wen2,
Lei Xie1#,
1Zhejiang University, 2Bytedance
*Equal Contribution. # Corresponding Author.
🌍 Overview
We introduce OpenVE-3M, an open-source, large-scale, and high-quality dataset for instruction-based video editing. The OpenVE-3M dataset includes eight major… See the full description on the dataset page: https://huggingface.co/datasets/Lewandofski/OpenVE-3M.Kai0
KAI0
TODO
The advantage label will be coming soon.
Contents
About the Dataset
Load the Dataset
Download the Dataset
Dataset Structure
Folder hierarchy
Details
License and Citation
About the Dataset
~134 hours real world scenarios
Main Tasks
Task_A
Single task
Initial state: T-shirts are randomly tossed onto the table, presenting random crumpled configurations
Manipulation task: Operate… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab-org/Kai0.D2E-Original
D2E-Original
Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation
This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 273.4 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents.
What's included:
Video + Audio: H.264 encoded at FHD/QHD 60fps with game audio.
Input events: Keyboard… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/D2E-Original.cobotmagic_Sim_drawer_open_placeEgoNormia
EgoNormia: Benchmarking Physical-Social Norm Understanding
MohammadHossein Rezaei*,
Yicheng Fu*,
Phil Cuvin*,
Caleb Ziems,
Yanzhe Zhang,
Hao Zhu,
Diyi Yang,
🌎Website |
🤗 Dataset |
📄 arXiv |
📄 HF Paper
EgoNormia
EgoNormia is a challenging QA benchmark that tests VLMs' ability to reason over norms in context.
The datset consists of 1,853 physically grounded egocentric
interaction clips from Ego4D… See the full description on the dataset page: https://huggingface.co/datasets/open-social-world/EgoNormia.prelinger-archives-open
Prelinger Archives Open License Videos
A collection of historical films from the Prelinger Archives on the Internet Archive, filtered to include only videos with open licenses (Public Domain, CC0, CC BY, CC BY-SA).
Dataset Description
The Prelinger Archives is a collection of over 17,000 advertising, educational, industrial, and amateur films. This dataset contains the subset of videos that are available under open licenses, making them freely usable for research… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/prelinger-archives-open.opencode_seed2.1_expert_skill_round_00_20260712AccidentBench
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
Website
·
Code
·
Leaderboard
·
Dataset
·
Dataset-Zip
·
Issue
Project Homepage:
https://accident-bench.github.io/
About the Dataset:
This benchmark includes approximately 2,000 videos and 19,000 human-annotated question-answer pairs, covering a wide range of reasoning tasks (as shown in Figure 1). We… See the full description on the dataset page: https://huggingface.co/datasets/Open-Space-Reasoning/AccidentBench.RMC-AIDA-L_pull_open_bag
RMC-AIDA-L_pull_open_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pull
zip
up
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_pull_open_bag.RoboCasa365aloha_static_cups_openThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_static_cups_open.LIBERO
LIBERO native-action EEF10 (LeRobot v3.0)
This local dataset was converted from lerobot/libero at revision
a1aaacb7f6cd6ee5fb43120f673cebb0cfea7dd4 (local source: /mnt/data/wangyuran/libero-lerobot). Its original parquet packing,
episode offsets, tasks, 10 FPS timeline, and videos are preserved.
Both primary robot columns use a 10-D interface:
observation.state: achieved [xyz3, rot6d6, gripper_open_scale1]
action: native normalized LIBERO [delta_xyz3, delta_rot6d6… See the full description on the dataset page: https://huggingface.co/datasets/OpenWAM/LIBERO.openarm-packingbench-v2-rawGalaxea-Open-World-Dataset_10K_20260123example_datasetDataset preview available at: https://huggingface.co/spaces/open-world-agents/visualize_dataset
OpenS2V-Eval
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
If you like our project, please give us a star ⭐ on GitHub for the latest update.
We release the high-quality OpenS2V-5M subset. It’s not just 0.3M samples — we applied filtering across the entire 5M data. You can click here for more details, and click here to download.
Regarding how to use OpenS2V-5M during the training phase, we provide a demo dataloader here. Alternatively, you can… See the full description on the dataset page: https://huggingface.co/datasets/BestWishYsh/OpenS2V-Eval.AgiBotWorld-Beta_G1_task_428_Open_the_drawer_and_store_items
agibot_task_428
This dataset converts the AgiBot format uniformly into LeRobot V3.0.
Dataset Statistics
robot_name: G1
end_effector: 夹爪
task: 打开抽屉,存放物品
total_episodes: 1782
total_tasks: 1
size: 138G
Dataset Structure
├── data
│ └── chunk-xxx
│ ├── file-xxx.parquet
├── meta
│ ├── episodes
│ │ └── chunk-xxx
│ │ └── file-xxx.parquet
│ ├── info.json
│ ├── stats.json
│ └── tasks.parquet
└── videos
├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-DataCube/AgiBotWorld-Beta_G1_task_428_Open_the_drawer_and_store_items.
