datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cad-1000-hours
CAD-1K Open v2 - 1,018.1229 Hours
509 end-to-end, single-display Windows CAD task recordings across seven CAD software families.
Each task contains:
task_desc.json - task prompt, application, reference-input paths, and expected deliverables
input_files/ - reference inputs named input.ext or input_N.ext
output_files/ - submitted CAD deliverables and supplemental outputs named output.ext or output_N.ext
rubrics.json - task-specific evaluation criteria
task_overview.pdf - review… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-1000-hours.xperience-10m
⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified.
Interactive Intelligence from Human Xperience
Xperience-10M
Dataset Summary
Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.L2DTL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot goes to driving school
90+ TeraBytes of multimodal data (5000+ hours of driving) from 30 cities in Germany
6x surrounding HD cameras and complete vehicle state: Speed/Heading/GPS/IMU
Continuous: Gas/Brake/Steering and discrete actions: Gear/Turn Signals
Environment state: Lane count, Road type (highway|residential), Road surface (asphalt, cobbled, sett), Max speed limit.… See the full description on the dataset page: https://huggingface.co/datasets/yaak-ai/L2D.cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.computer-use-large
Computer Use Large
A large-scale dataset of 48,478 screen recording videos (~12,300 hours) of professional software being used, sourced from the internet. All videos have been trimmed to remove non-screen-recording content (intros, outros, talking heads, transitions) and audio has been stripped.
Dataset Summary
Category
Videos
Hours
AutoCAD
10,059
2,149
Blender
11,493
3,624
Excel
8,111
2,002
Photoshop
10,704
2,060
Salesforce
7,807
2,336
VS… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/computer-use-large.tracker-pov
Eidon Tracker POV
1,274 hours of egocentric video paired with 7-point IMU arm tracking, recorded during ordinary household work.
Contributors wore a head-mounted camera and a seven-sensor IMU harness while doing real chores in their own homes: laundry, cleaning, dishes, cooking. Each recording pairs first-person video with 24 Hz orientation data for both hands, both forearms, both upper arms, and the chest.
This is a complete, final release. Eidon AI (Solidic Labs Inc) has wound… See the full description on the dataset page: https://huggingface.co/datasets/eidon-ai/tracker-pov.nexar_collision_prediction
Nexar Collision Prediction Dataset
This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle.
Dataset
The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.UltraVideo
UltraVideo: High-Quality UHD 4K Video Dataset
🤓 Project | 📑 Paper | 🤗 Hugging Face (UltraVideo Dataset)) | 🤗 Hugging Face (UltraVideo-Long Dataset)) | 🤗 Hugging Face (UltraWan-1K/4K Weights)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
🎋 Click below image to watch the 4K demo video.
🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.epic-kitchens-100-clips
EPIC-KITCHENS-100 Extracted Clips
About
Dataset of 37455 video clips (24GB) extracted from videos in the EPIC-KITCHENS-100 dataset,
more precisely the extension part not contained in EPIC-KITCHENS-55. For details,
see https://www.lightly.ai/product-updates/epickitchens-100-in-lightlystudio.
The clips folder contains one video for every narration from action annotations stored
in {participant_id}/{narration_id}.mp4. The videos have been downscaled an compressed for easier… See the full description on the dataset page: https://huggingface.co/datasets/lightly-ai/epic-kitchens-100-clips.groot_n1.7_inference_on_diff_data
GROOT Inference Analysis Log
Evaluation records for a GR00T policy trained on the task
"pick octopus and place inside brown basket", run on a Unitree G1 at 20 Hz
with the ego_view stereo camera.
Six training checkpoints (50, 100, 150, 200, 250, 300 demonstration episodes) were each
evaluated on 50 inference episodes. Every episode is recorded here with video,
per-tick state/action logs and run metadata.
Success rate
Checkpoint (training episodes)
Success… See the full description on the dataset page: https://huggingface.co/datasets/rahul-ai-01/groot_n1.7_inference_on_diff_data.airoa-moma-5k
AIRoA MoMa 5k
AIRoA MoMa 5k is a large-scale, task-structured dataset of real-robot mobile manipulation collected by teleoperating Toyota Human Support Robots (HSRs). The public release contains 1,184,259 successful Primitive-Action (PA) episodes, 180,905,084 frames, and 5,025 recorded hours from 44 physical robots at five collection sites.
Each PA remains independently addressable for policy training, while execution-level metadata preserve the Short-Horizon Task (SHT) in which… See the full description on the dataset page: https://huggingface.co/datasets/airoa-org/airoa-moma-5k.AIGVDBenchhumanplus-1000
HumanPlus-1000
HumanPlus-1000 is a large-scale multimodal human behavior dataset
designed for learning and modeling human perception, motion, and interaction
in real-world environments.
It captures synchronized egocentric visual observations, full-body motion,
hand motion, camera motion, and 3D environment information, with the goal of
providing paired human perception–action data for embodied intelligence,
human motion modeling, and robotics.
Current Release: This repository… See the full description on the dataset page: https://huggingface.co/datasets/humanplus-ai/humanplus-1000.UltraVideo-Long
UltraVideo: High-Quality UHD 4K Video Dataset
🤓 Project | 📑 Paper | 🤗 Hugging Face (UltraVideo Dataset)) | 🤗 Hugging Face (UltraVideo-Long Dataset)) | 🤗 Hugging Face (UltraWan-1K/4K Weights)
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
🎋 Click below image to watch the 4K demo video.
🤓 First open-sourced UHD-4K/8K video datasets with comprehensive structured (10 types) captions.🤓 Native 1K/4K videos generation by UltraWan.… See the full description on the dataset page: https://huggingface.co/datasets/APRIL-AIGC/UltraVideo-Long.vifailback-dataset-lerobot
ViFailback Dataset — LeRobot
This repository is a LeRobot v2.1 conversion of the trajectory portion of sii-rhos-ai/ViFailback-Dataset, introduced in the CVPR 2026 paper Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols.
ViFailback contains real-world ALOHA dual-arm manipulation trajectories designed for studying failure diagnosis, failure localization, corrective guidance, recovery, and learning… See the full description on the dataset page: https://huggingface.co/datasets/sii-rhos-ai/vifailback-dataset-lerobot.SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.wearable-ai
EgoWearBench Dataset (ECCV 2026)
Part of the Wearable AI Workshop at ECCV 2026.
A benchmark of egocentric (first-person, head-mounted wearable camera) videos paired with three complementary video question-answering tasks for evaluating wearable-AI assistants on real-world everyday activity videos.
▶ Baseline code & evaluation scripts: see starter_kit/README.md. The starter kit ships inside this repo, so git clone gives you the code and the data together.
Tasks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/wearable-ai.gm100-cobotmagic-lerobot
Dataset Card for GM-100 Cobot Magic Part (Lerobot2.1 Format)
This is a part of the Great March 100 (GM-100) Project.
Raw teleoperation data collected using the Agilex Cobot Magic robot. This dataset is stored in Lerobot2.1 format.
Notice
This is a preview version of this dataset. A few task prompts may contain minor issues, which we will address in the upcoming full release. We welcome any feedback or suggestions you may have.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/rhos-ai/gm100-cobotmagic-lerobot.GR1_robot
GR1 Robot Teleoperation Dataset
A large-scale humanoid robot teleoperation dataset collected using the Fourier GR-1 (GR1T1) robot, part of the GR00T robotics initiative.
This is a true subset of nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.
Dataset Summary
Episodes: 22,209
Total frames: 6,362,293
Total tasks: 46,669
Videos: 22,209
Frequency: 20 FPS
Robot: GR1T1 (Fourier GR-1 humanoid)
The dataset contains egocentric video observations (1280x800) paired with full-body… See the full description on the dataset page: https://huggingface.co/datasets/Physis-AI/GR1_robot.pretrain_aiworker_bg2_lance
rllab-postech/pretrain_aiworker_bg2_lance
Merged 19D AI Worker/BG2 pretraining dataset in RLLAB published Lance layout.
Tables
Table
Purpose
data/episodes.lance
Published episode table, one row per episode, no video blob columns.
data/train_episodes.lance
Training trajectory table named by manifest.json.primary_training_table; no video blob columns.
data/frames.lance
Frame-level QA/index table with remapped global frame indices.
data/videos.lance… See the full description on the dataset page: https://huggingface.co/datasets/rllab-postech/pretrain_aiworker_bg2_lance.ReactNet
ResponseNet
ResponseNet is a large-scale dyadic video dataset designed for Online Multimodal Conversational Response Generation (OMCRG). It fills the gap left by existing datasets by providing high-resolution, split-screen recordings of both speaker and listener, separate audio channels, and word‑level textual annotations for both participants.
Paper
If you use this dataset, please cite:
ResponseNet: A High‑Resolution Dyadic Video Dataset for Online Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/awakening-ai/ReactNet.AirScape-Dataset
[ACM MM'25] AirScape: An Aerial Generative World Model with Motion Controllability
This repository contains the dataset introduced in the paper, consisting of two parts: 11k+ motion intention prompts and corresponding video clips.
Arxiv: https://arxiv.org/pdf/2507.08885
Project: https://embodiedcity.github.io/AirScape/
Code: https://github.com/EmbodiedCity/AirScape.code
Dataset Description
This dataset is proposed for training and testing of aerial world models.… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/AirScape-Dataset.physical-ai-bench-step-30000-videos
PhysicalAIBench Step 30000 Model Comparison
This repository contains two filename-aligned sets of 5,220 MP4 outputs from
step_30000 evaluation runs. The files are presented through the
companion PhysicalAI Video Gallery.
Model sets
Directory
Model
Files
Bytes
videos/
DC-AE v0.2, Cosmos encoder + causal decoder
5,220
4,244,965,541
videos_wan22_vae/
Wan 2.2 VAE, phase 3 PDX
5,220
4,818,900,749
The two directories have an exact 1:1 basename match.… See the full description on the dataset page: https://huggingface.co/datasets/spongy/physical-ai-bench-step-30000-videos.RMC-AIDA-L_glasses_storage
RMC-AIDA-L_glasses_storage
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
fold
close
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_glasses_storage.RMC-AIDA-L_desktop_organization
RMC-AIDA-L_desktop_organization
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_desktop_organization.RMC-AIDA-L_pull_open_bag
RMC-AIDA-L_pull_open_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pull
zip
up
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_pull_open_bag.RMC-AIDA-L_food_packaging
RMC-AIDA-L_food_packaging
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
pull
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_food_packaging.AIRBOT_MMK2_building_block_storage
AIRBOT_MMK2_building_block_storage
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_building_block_storage.RMC-AIDA-L_fold_shirt
RMC-AIDA-L_fold_shirt
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
fold
place
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_fold_shirt.
