datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.fractal20220817_data_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 87212,
"total_frames": 3786400,
"total_tasks": 599,
"total_videos": 87212,
"total_chunks": 88,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:87212"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.LLaVA-OneVision-2-Data
LLaVA-OneVision-2-Data
Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training.
At a Glance
The dataset is split across two Hugging Face repositories because of its size:
Repository
What it contains
Part 1 (this repository)
~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.HiFi-UMI-2K
HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data
2,000 hours released · 6 synchronized camera views · 480+ scenes · 3 mm pose accuracy · <40 µs synchronization
🌐 Project Website |
📦 Dataset |
📄 Paper: arXiv:2607.25895
Examples from the HiFi-UMI corpus. Click the image to play the video.
📚 Introduction
HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations.… See the full description on the dataset page: https://huggingface.co/datasets/simple-world-lab/HiFi-UMI-2K.2026-challenge-demos
BEHAVIOR-1K 2026 Challenge Demos
This dataset contains BEHAVIOR-1K 2026 challenge demonstration trajectories in LeRobotDataset v3 format.
Dataset Statistics
Tasks: 100
Episodes: 20,000
Frames: 210,916,774
Size: approximately 3.0 TB
Data shards: 955 Parquet files
Video files: 17,093 MP4 files
Video features: 6
Format
The repository follows the LeRobotDataset v3 layout:
meta/info.json: dataset schema and path templates
meta/stats.json: feature… See the full description on the dataset page: https://huggingface.co/datasets/behavior-1k/2026-challenge-demos.2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/behavior-1k/2025-challenge-demos.bridgev2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "WidowX",
"total_episodes": 53192,
"total_frames": 1999410,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Qu3tzal/bridgev2.MolmoAct2-BimanualYAM-DatasetThis dataset was created using LeRobot.
MolmoAct2-BimanualYAM Dataset
This repository is the merged ckpt / merged LeRobot dataset artifact for the MolmoAct2-BimanualYAM Dataset, a large-scale collection of bimanual robot manipulation demonstrations collected for MolmoAct2. Across the full collection, MolmoAct2-BimanualYAM contains more than 720 hours of training demonstrations spanning diverse tabletop manipulation tasks.
Language Annotations
This dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-BimanualYAM-Dataset.2deform360
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Project Page | Paper | GitHub Repository
Deform360 is a massive multi-view visuotactile dataset for deformable-object research, featuring 198 daily-life objects, 1,980 interaction sequences, and over 215 hours of observations from 41 surround-view cameras and bimanual tactile grippers to capture both global motion and contact-induced local deformations.
Installation
To load and… See the full description on the dataset page: https://huggingface.co/datasets/liuyibing/2deform360.DH-FaceVid-1KGOKU-2M
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
GOKU-2M is a large-scale, unified instruction-based video-editing dataset covering 10 editing tasks. Each sample provides a source video, an edited target video, and one or more natural-language instructions describing the edit.
📦 Repositories
⚠️ Because a single Hugging Face account has a free storage quota of about 8.7 TB, the dataset is split across two… See the full description on the dataset page: https://huggingface.co/datasets/bigfacing/GOKU-2M.VBench-2.0_sampled_videos
Sample Videos of VBench-2.0
This dataset is used in the paper:👉 arXiv:2503.21755
DataDump_07-17-2026
RoboArena Dataset Snapshot — 2026-07-17
This dataset contains autonomous policy rollouts, task-success
scores, policy-preference annotations, and related metadata collected by the
RoboArena benchmark from its inception through July 17, 2026.
Snapshot statistics
Evaluation sessions: 3,883
Policy episodes: 10,783
MP4 videos: 27,148
NPZ proprioception/action files: 10,783
Session metadata files: 3,883
Data layout
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/RoboArena/DataDump_07-17-2026.repo3InternData-fractal20220817_dataCASTLE2024
What is CASTLE?
The CASTLE dataset is a large-scale, multimodal dataset designed for advancing research in lifelogging, human activity recognition, and multimodal retrieval. It provides a rich collection of time-aligned sensor and video data for analysis and benchmarking. See the Paper (or its arXiv pre-print) for more details.
You can check our website for more details.
Characteristics
Captured over four days in a controlled environment
10 participants engaged… See the full description on the dataset page: https://huggingface.co/datasets/CASTLE-Dataset/CASTLE2024.h3-fewstep-benchmark-20260920
H3 few-step benchmark
Fifteen few-step checkpoints for the MiniMax-H3 video+audio model, each run on the
same 48 English prompts × 2 seeds (0 and 1) × 3 step counts (4, 8, 32) =
288 videos per checkpoint, 4,320 videos in total. Conditioning is text only
(no input image, no camera trajectory). Every video is 1344×768, 5 seconds
(120 frames at 24 fps), with synthesized audio unless noted.
All checkpoints are adapters or distilled variants of MiniMaxAI/MiniMax-H3, except the… See the full description on the dataset page: https://huggingface.co/datasets/hffordata/h3-fewstep-benchmark-20260920.bridge_v2_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "widowx",
"total_episodes": 53192,
"total_frames": 1999410,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jesbu1/bridge_v2_lerobot.egoscaler-v2
EgoScalerV2 Dataset
This dataset accompanies our work on Developing Vision-Language-Action Model from Egocentric Videos. It provides 6DoF object trajectories paired with egocentric visual observations and natural-language action descriptions, formatted in the LeRobot v2.0 schema so it can be consumed directly by LeRobot-compatible pipelines.
🌐 Project page: https://biscue5.github.io/egovla-project-page/
📄 Paper: Developing Vision-Language-Action Model from Egocentric Videos… See the full description on the dataset page: https://huggingface.co/datasets/Biscue5/egoscaler-v2.Video-MME-v2
🔥 News
2026.06.11 Videos re-encoded to H265, maintaining consistent evaluation scores. Fixed 2 incorrect MP4s & 3 mismatched URLs. Original data preserved in the original branch.
2026.05.22 Task types are now available for Q1-Q3 in coherence (logic) groups.
🤗 About This Repo
This repository contains annotation data for "Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding". It mainly consists of three… See the full description on the dataset page: https://huggingface.co/datasets/MME-Benchmarks/Video-MME-v2.behavior_224_rgbThis is the compressed version of the original BEHAVIOR dataset
It contains only RGB videos compressed to 224x224 as well as actions, annotations, and metadata files.
Depth and segmentation data are removed.
The dataset is just ~260GB, which makes it easier to use than the original one if you don't need all the data.
We used this dataset for our 1st place solution in the NeurIPS 2025 BEHAVIOR Challenge. Code, tech report.
Citation
@article{li2024behavior,
title={Behavior-1k:… See the full description on the dataset page: https://huggingface.co/datasets/IliaLarchenko/behavior_224_rgb.2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/elonelonelon/2025-challenge-demos.behavior-1k_2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dario-Shit4/behavior-1k_2025-challenge-demos.robocasa_22_tasksGOAI-2026OctoSense
New to OctoSense? A getting-started Colab notebook walks through loading a sequence and using each modality. Click Open in Colab above to run it in your browser, no setup required.
OctoSense is a time-synchronized, calibrated, multi-sensor dataset spanning multiple platforms, all sharing the same sensor rig. The bulk of the data is large-scale driving dataset: 371 sequences · 59 hrs · 2,474 km · 8.43 TB of urban, suburban, and rural driving… See the full description on the dataset page: https://huggingface.co/datasets/anthonytec2/OctoSense.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.GOKU-2M
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing
GOKU-2M is a large-scale, unified instruction-based video-editing dataset covering 10 editing tasks. Each sample provides a source video, an edited target video, and one or more natural-language instructions describing the edit.
📦 Repositories
⚠️ Because a single Hugging Face account has a free storage quota of about 8.7 TB, the dataset is split across two… See the full description on the dataset page: https://huggingface.co/datasets/Goku-2M/GOKU-2M.first-impressions-v2
Dataset Card for First Impressions V2
The first impressions data set, comprises 10000 clips (average duration 15s) extracted from more than 3,000 different YouTube high-definition (HD) videos of people facing and speaking in English to a camera. The videos are split into training, validation and test sets with a 3:1:1 ratio. People in videos show different gender, age, nationality, and ethnicity.
Videos are labeled with personality traits variables. Amazon Mechanical Turk (AMT) was… See the full description on the dataset page: https://huggingface.co/datasets/yeray142/first-impressions-v2.
