datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bridge_orig_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "widowx",
"total_episodes": 53192,
"total_frames": 1893026,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot.MotionBench
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset] [💻 GitHub] [🏆 Leaderboard] [🏆 HF Leaderboard]
MotionBench is a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. It evaluates models' motion-level perception through six primary categories of motion-oriented question types and includes data… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/MotionBench.scenetok_originairoa-moma-5k
AIRoA MoMa 5k
AIRoA MoMa 5k is a large-scale, task-structured dataset of real-robot mobile manipulation collected by teleoperating Toyota Human Support Robots (HSRs). The public release contains 1,184,259 successful Primitive-Action (PA) episodes, 180,905,084 frames, and 5,025 recorded hours from 44 physical robots at five collection sites.
Each PA remains independently addressable for policy training, while execution-level metadata preserve the Short-Horizon Task (SHT) in which… See the full description on the dataset page: https://huggingface.co/datasets/airoa-org/airoa-moma-5k.bridge_orig_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "widowx",
"total_episodes": 53192,
"total_frames": 1893026,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ryanqian1994/bridge_orig_lerobot.Robotic_Origami_Challenge
Robotic Origami Challenge: Fold Plane Demonstrations
Real-world LeRobot demonstrations for dexterous paper-airplane folding.
Overview
Robotic Origami Challenge: Fold Plane Demonstrations is a real-world teleoperation dataset for folding a paper airplane with a bimanual dexterous robot system. It is released by Sharpa in a LeRobot-compatible format for the Robotic Origami Challenge community.
Origami is a demanding benchmark for embodied AI:… See the full description on the dataset page: https://huggingface.co/datasets/SharpaIT/Robotic_Origami_Challenge.brand-assetsdataset
YouTube Video Downloads
This dataset contains locally downloaded YouTube video files from /data/youtube_downloads.
Source files scanned: 88549
Uploaded unique filenames: 87367
Duplicate filenames resolved: 1182
Selected video bytes: 7103828648301
Manifest rows: 87367
Files are stored under videos/. Metadata is stored under metadata/.
metadata/manifest.csv includes these columns:
target_path, filename, video_id, title, duration_seconds, duration, width, height, exact_width… See the full description on the dataset page: https://huggingface.co/datasets/Orannue/dataset.Kai0
KAI0
TODO
The advantage label will be coming soon.
Contents
About the Dataset
Load the Dataset
Download the Dataset
Dataset Structure
Folder hierarchy
Details
License and Citation
About the Dataset
~134 hours real world scenarios
Main Tasks
Task_A
Single task
Initial state: T-shirts are randomly tossed onto the table, presenting random crumpled configurations
Manipulation task: Operate… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab-org/Kai0.D2E-Original
D2E-Original
Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation
This is the dataset for D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI. 273.4 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open-world, sandbox, and more), for training vision-action models and game agents.
What's included:
Video + Audio: H.264 encoded at FHD/QHD 60fps with game audio.
Input events: Keyboard… See the full description on the dataset page: https://huggingface.co/datasets/open-world-agents/D2E-Original.data_verticalAgilex_Cobot_Magic_basket_storage_orange
Agilex_Cobot_Magic_basket_storage_orange
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_basket_storage_orange.umi-okra-dex1
UMI Okra Grasping Dataset — Lab Subset (Unitree Dex1-1)
Universal Manipulation Interface (UMI) hand-held teleoperation data for an okra-fruit
grasping / harvesting task. Recorded to train a Diffusion Policy (with a comparison
ACT track) deployed on a Unitree G1.
This is the indoor lab subset. Every session here was recorded in a lab mock okra
field (artificial foliage, white-walled room). The outdoor sessions from the original
collection are not included — see Scope and… See the full description on the dataset page: https://huggingface.co/datasets/Orboh/umi-okra-dex1.Cobot_Magic_desktop_organization
Cobot_Magic_desktop_organization
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_desktop_organization.G1_Dex1_OrganizeToolsRMC-AIDA-L_desktop_organization
RMC-AIDA-L_desktop_organization
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_desktop_organization.DeepSea-MOT
DeepSea MOT
DeepSea MOT is a benchmark dataset for multi-object tracking on deep-sea video.
Dataset Description
DeepSea MOT consists of 4 video sequences (2 midwater, 2 benthic) with a total of 2,400 frames and 57,376 annotated objects comprising 188 tracks. The videos were captured by the Monterey Bay Aquarium Research Institute (MBARI) using remotely operated vehicles (ROVs) Doc Ricketts and Ventana in deep-sea environments, showcasing a variety of marine species and… See the full description on the dataset page: https://huggingface.co/datasets/MBARI-org/DeepSea-MOT.LVBench
LVBench: An Extreme Long Video Understanding Benchmark
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset][🏆 Leaderboard]
LVBench is a benchmark designed to evaluate and enhance the capabilities of multimodal models in understanding and
extracting information from long videos up to two hours in duration.
🔥 News
2024.06.11 🌟 We released LVBench, a new benchmark for long video understanding!
👀 Introduce to LVBench
LVBench is a benchmark designed to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LVBench.G1_Dex1_Organize_ToolsThis dataset was created using LeRobot.
Due to the inability to precisely describe spatial positions, adjust the scene to closely match the first frame of the dataset after installing the hardware as specified in Part 5 of AVP Teleoperation Documentation.
Data collection is not completed in a single session, and variations between data entries exist. Ensure these variations are accounted for during model training.
Dataset Structure
meta/info.json:
{
"codebase_version":… See the full description on the dataset page: https://huggingface.co/datasets/unitreerobotics/G1_Dex1_Organize_Tools.heico-focus-vqa
HeiCo-FOCUS (Beta release)
A clinically grounded dataset for long-context video understanding in minimally invasive surgery.
📄 Paper • 🤗 Dataset • 💻 Code • 🏆 Challenge • ⚖️ CC BY-NC-SA 4.0
[!NOTE]
Until 31 July 2026, we will employ a restricted post-release review period. During this time, we kindly ask the community to provide feedback to help us increase quality control and make continuous improvements.… See the full description on the dataset page: https://huggingface.co/datasets/orena-dkfz/heico-focus-vqa.dvla-multi10x5-orig-25fpsgame-recordings-v3
OriginLab Game Recordings v0.3.0
Human gameplay captured in-engine under per-title licenses at 1080p / 60 FPS CFR on one shared frame clock: every stream starts at frame 0 and frame k matches frame k across pre-HUD and post-HUD RGB, surface normals, metric depth, audio, camera telemetry, keyboard and mouse inputs, in-engine action events and game state, and world telemetry — plus per-frame training tables.
Watch full playable previews of every modality, side by side and in… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-recordings-v3.lapchole-focus-vqa
LapChole-FOCUS-VQA
A clinically grounded benchmark for long-context video understanding in minimally invasive surgery.
💻 Code • 🏆 Challenge • ⚖️ Data Usage Agreement
[!IMPORTANT]
🔒 This is a gated dataset
Access is granted only to participants of the ORena FOCUS Challenge and is subject to manual review. To be approved you must:
Have a Hugging Face account and be logged in — downloads are only enabled for registered, authenticated… See the full description on the dataset page: https://huggingface.co/datasets/orena-dkfz/lapchole-focus-vqa.OraRL-Data
OraRL-Data
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Video-ORA-9B] [💻 Code]
We release OraRL-Data, the official evaluation suite for Video-ORA and OraRL.
It packages the canonical annotations and referenced raw media used by the OraRL evaluation suite: 109,374 examples across 16 benchmark configs and 29 splits, with 518.9 GiB of manifested files. The complete evaluation release lives under OraRL-eval-data/, leaving room for the separate OraRL training release in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/OraRL/OraRL-Data.RMC-AIDA-L_organise_the_document_bag
RMC-AIDA-L_organise_the_document_bag
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
pull
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_organise_the_document_bag.Ordering_ConstrainedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Ordering_Constrained.RMC-AIDA-L_basket_storage_orange
RMC-AIDA-L_basket_storage_orange
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: realman_rmc_aidal
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/RMC-AIDA-L_basket_storage_orange.yubi-corl2026-umi-arena
YUBI — The UMI Arena @ CoRL 2026
This is the initial competition data release for The UMI Arena,
a workshop at CoRL 2026 (November 12, 2026 — Austin, Texas, USA) organized by the
AI Robot Association (AIRoA). The release is a large-scale dataset of real-robot
bimanual manipulation
collected by teleoperating the Yubi robot — a bimanual platform with
two-finger grippers and per-fingertip pose tracking — through a Meta Quest
headset, hand-held controllers, and a foot pedal. The… See the full description on the dataset page: https://huggingface.co/datasets/airoa-org/yubi-corl2026-umi-arena.g1_desktop_organize_table_correction
g1_desktop_organize_table_correction
Bimanual tabletop teleoperation on a Unitree G1 with Inspire dexterous hands,
in LeRobot v2.1 format.
Task — Follow the human's corrective gesture and hand the indicated object to the human.
Episodes
192
Frames
67405 (37.4 min @ 30 fps)
Episode length
268–519 frames (median 346)
State / action
26-D / 26-D
Cameras
2 × 640×480
State and action
Both vectors are 26-D and share the same layout:
0– 6… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/g1_desktop_organize_table_correction.SP_OrderingThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 24037,
"total_tasks": 10,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/SP_Ordering.
