datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicTran38K
PhysicTran38K
We provide 38K video-based dataset for physics-aware image editing, by casting editing as physical state transitions. To know the details of the dataset creation pipeline, please refer to Section 3.2 of [1].
Dataset Download and Loading
huggingface-cli download diffusion-cot/PhysicTran38K --local-dir /path/to/data_path
Then please refer to our training code for data loading.
✍️ Citation
If you find our work inspiring or use our codebase in… See the full description on the dataset page: https://huggingface.co/datasets/metazlb/PhysicTran38K.MetaFine_CoRL26
CoRL demos
Training data for the MetaFine π0 baseline (T1–T5).
Overview
Task
Variant tree
Mixed episodes
Mixed frames
Cameras
State
FPS
Control / obs
T1 grasp_part
cap / body / mixed
200
30,670
base + hand, 512×512, FOV ≈ 70°
9-DoF
30
pd_joint_delta_pos / rgb
T2 grasp_move_mug
left / right / forward / mixed
300
39,395
base + hand, 512×512, FOV ≈ 70°
9-DoF
30
pd_joint_delta_pos / rgb
T3 toggle_switch_table
red / blue / mixed
200
20,567
base + hand… See the full description on the dataset page: https://huggingface.co/datasets/hiangx/MetaFine_CoRL26.Evo1_MetaWorld_Datasetmeta-archivemeta-libero-resultsmetal_part_sort_v10_plus_extra_20260702
metal_part_sort_v10_plus_extra_20260702
Merged LeRobot v2.1 dataset for Unitree G1 + Inspire DFX metal part sorting.
Base dataset: PID0930/metal_part_sort_v10 @ 17fe43673ca30e840b1baa6e05c1f34b7dd43b91
Extra dataset: PID0930/metal_part_sort_v10_extra_20260701 @ 09c5a9113830443f99a8d91c11fd5754b238fc6a
Episodes: 329
Frames: 110966
Cameras: external, left wrist, right wrist, plus left_high/head slot retained for compatibility
GR00T training uses external as the logical head… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/metal_part_sort_v10_plus_extra_20260702.BenchCheck-MetaBenchmark
BenchCheck meta-benchmark (v6)
840 multiple-choice video questions from 76 public video benchmarks, one item per video,
four capability groups x 210 (Perception, Temporal, Spatial / physical, Reasoning / knowledge). The set is the
difficulty-first census of the BenchCheck screened pool: items that none of the cheap attackers of the pool screening
(text-only, single frame, 32-frame 2B model, options-only, shuffled frames) could solve, ranked by the worst-case attacker
percentile… See the full description on the dataset page: https://huggingface.co/datasets/GMLRVigil/BenchCheck-MetaBenchmark.loop_metal_20260902_132356This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
7
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/loop_metal_20260902_132356.G1edu-u3_pick_metal_bowl_aa
G1edu-u3_pick_metal_bowl_aa
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: G1edu-u3
| Codebase Version: v2.1
End-Effector Type: three_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_pick_metal_bowl_aa.G1edu-u3_place_metal_bowl_ae
G1edu-u3_place_metal_bowl_ae
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: G1edu-u3
| Codebase Version: v2.1
End-Effector Type: three_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
pick
place
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_place_metal_bowl_ae.G1edu-u3_place_metal_bowl_af
G1edu-u3_place_metal_bowl_af
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: G1edu-u3
| Codebase Version: v2.1
End-Effector Type: three_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
pick
place
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_place_metal_bowl_af.G1edu-u3_pick_metal_bowl_ab
G1edu-u3_pick_metal_bowl_ab
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: G1edu-u3
| Codebase Version: v2.1
End-Effector Type: three_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
📊 Dataset Statistics
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_pick_metal_bowl_ab.pickup-carrot-remove-parquet-metadata-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata-2.placement_yellow_dot_red_arrow_hdf5_metadata_100fps
Placement Yellow-Dot Red-Arrow HDF5 Metadata Dataset
This is a LeRobot-format full dataset export derived from local HDF5 placement data.
Dataset size:
episodes: 197
frames: 66397
videos: 591
export fps: 100
frame stride from 100 Hz source: 1
Instruction:
move the connector to the yellow dot and match the red arrow orientation
The overhead video is rendered from the raw HDF5 overhead camera frame with an anti-aliased red arrow and yellow source dot plotted from HDF5… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/placement_yellow_dot_red_arrow_hdf5_metadata_100fps.metaworldThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "metaworld",
"total_episodes": 2500,
"total_frames": 204806,
"total_tasks": 49,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 80,
"splits": {
"train": "0:2500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jadechoghari/metaworld.placement_yellow_dot_red_arrow_hdf5_metadata
Placement Yellow-Dot Red-Arrow HDF5 Metadata Dataset
This is a LeRobot-format full dataset export derived from local HDF5 placement data.
Dataset size:
episodes: 197
frames: 66397
videos: 591
Instruction:
move the connector to the yellow dot and match the red arrow orientation
The overhead video is rendered from the raw HDF5 overhead camera frame with an anti-aliased red arrow and yellow source dot plotted from HDF5 metadata:
placement/episode_target_image_xy_px… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/placement_yellow_dot_red_arrow_hdf5_metadata.Carrot_data
Carrot_data
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
Banana_data
Banana_data
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
placement_yellow_dot_red_arrow_hdf5_metadata_50fps_subsampled
Placement Yellow-Dot Red-Arrow HDF5 Metadata Dataset
This is a LeRobot-format full dataset export derived from local HDF5 placement data.
Dataset size:
episodes: 197
frames: 33247
videos: 591
export fps: 50
frame stride from 100 Hz source: 2
Instruction:
move the connector to the yellow dot and match the red arrow orientation
The overhead video is rendered from the raw HDF5 overhead camera frame with an anti-aliased red arrow and yellow source dot plotted from HDF5 metadata:… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/placement_yellow_dot_red_arrow_hdf5_metadata_50fps_subsampled.pickup-carrot-remove-parquet-metadataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 21,
"total_frames": 9383,
"total_tasks": 1,
"total_videos": 84,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:21"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-remove-parquet-metadata.metaworld_mt50_lerobot21This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "metaworld",
"total_episodes": 2500,
"total_frames": 204806,
"total_tasks": 49,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 80,
"splits": {
"train": "0:2500"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/DorayakiLin/metaworld_mt50_lerobot21.metaworld_dataset_v2.1metaworld_mt50_25_08_22_lerobotv2.1metal_part_sort_v6
metal_part_sort_v6
LeRobot v2.1 dataset for the Unitree G1 + Inspire DFX metal part sorting task.
Episodes: 39
Frames: 32,049
FPS: 30
Raw snapshot: /home/y-kosugi/datasets/metal_part_sort_v6_all_raw_20260624/metal_part_sort_v6
Camera streams are stored as:
observation.images.cam_left_high: robot head camera, recorded for audit only
observation.images.cam_external: external camera
observation.images.cam_left_wrist: left wrist camera
observation.images.cam_right_wrist: right… See the full description on the dataset page: https://huggingface.co/datasets/PID0930/metal_part_sort_v6.metal_part_sort_v3
metal_part_sort_v3
LeRobot v2.1 dataset for Unitree G1 + Inspire DFX metal part sorting.
Episodes: 95
Frames: 46748
Cameras: head, left wrist, right wrist
State/action: 26D, with hand dimensions 14:26
Hand-state convention: observation.state[t, 14:26] = action[t-1, 14:26] for t > 0; frame 0 uses same-frame hand action.
Failed recordings moved to trash were excluded before conversion.
metal_part_sort_v10
metal_part_sort_v10
LeRobot v2.1 dataset recorded on Unitree G1 + Inspire DFX for metal part sorting.
Episodes: 309
Frames: 101594
Cameras: external, left wrist, right wrist, head slot mapped to external for GR00T compatibility
State/action: 26 dims (left arm 7, right arm 7, left hand 6, right hand 6)
C_plan_train_shard00_meta
Causal_Plan Dataset
A multimodal dataset for fine-tuning Vision-Language Models (VLMs). It processes egocentric video (Ego4D, EPIC-Kitchens) into structured causal plans, generates 462K multimodal QA pairs across 24 task types, and exports them for LoRA SFT of Qwen3-VL-8B-Instruct.
Quick Start
Prerequisites
huggingface-cli (install via pip install huggingface_hub)
~900 GB free disk space
A HuggingFace token with read access (huggingface-cli login)… See the full description on the dataset page: https://huggingface.co/datasets/Lululzz/C_plan_train_shard00_meta.metal_part_sort_v10_extra_20260701
metal_part_sort_v10_extra_20260701
Extra LeRobot v2.1 dataset recorded on Unitree G1 + Inspire DFX for metal part sorting.
Source raw episodes: episode_0311 through episode_0330 from local metal_part_sort_v10
Episodes: 20
Frames: 9372
Cameras: external, left wrist, right wrist, plus left_high/head slot retained for compatibility
GR00T training uses external as the logical head camera input
State/action: 26 dims (left arm 7, right arm 7, left hand 6, right hand 6)
so101_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 2,
"total_frames": 1791,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Metabolik/so101_test.metal_part_sort_v5
