datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
so101-egg-transport-act-v0
SO-101 raw egg transport: teleoperation and first ACT rollout
This dataset documents a single-arm Hiwonder SO-ARM101 task: pick up a raw egg
from a fixed pickup area and gently place it in an unheated frying pan.
Contents
Root dataset: 25 human teleoperation demonstrations in LeRobot v3 format.
evaluations/autonomous_success_001/ through
evaluations/autonomous_success_007/: seven saved autonomous-success
rollouts of an ACT policy trained from the 25… See the full description on the dataset page: https://huggingface.co/datasets/NOJIMA21/so101-egg-transport-act-v0.so101-leader-urdf
SO-101 leader URDF
Open so101_leader_new_calib.urdf with the adjacent assets/ directory intact. All mesh paths are relative; all mesh coordinates are in metres. This is a geometry/kinematics conversion of the supplied follower URDF, with a provisional trigger calibration.
Changes
Preserved the original base, shoulder, upper arm, lower arm, wrist, motor meshes, and the five arm joint origins, axes, limits and transmissions.
Replaced the fixed follower gripper body… See the full description on the dataset page: https://huggingface.co/datasets/cetiennec/so101-leader-urdf.so101_pick_cubeDexdata format SO-101 dataset
Dataset Structure
so101_pick_cube
├── videos
│ ├── so101_20260630_103330_filtered
│ │ ├── file-000.mp4_top.mp4
│ │ └── file-000.mp4_wrist.mp4
│ └── ...
└── jsonl
├── episode_00000.jsonl
├── episode_00001.jsonl
└── ...
Task Description
This dataset contains robot manipulation demonstrations for the task:
Task: Pick the cube and place it in the plate
Object variations:… See the full description on the dataset page: https://huggingface.co/datasets/Dexmal/so101_pick_cube.so101-sword-on-stand-clean32-bothtrim-v2
SO-101 sword-on-stand clean32, physically both-end trimmed
This is the corrected, portable LeRobot v2.1 derivative for the task:
Pick up the sword and place it on the sword stand.
Dataset contract
32 episodes / 6,974 frames / 30 Hz.
Train episodes: 0..29 (6,582 frames).
Validation episodes: 30, 31 (392 frames).
Fixed and wrist RGB cameras, 640x480, H.264, 30 FPS.
observation.state and stored action are calibrated 6D absolute joint positions.
Only episodes 0..29… See the full description on the dataset page: https://huggingface.co/datasets/WDong/so101-sword-on-stand-clean32-bothtrim-v2.SO101_test109
Model Card for act
Action Chunking with Transformers (ACT) is an imitation-learning method that predicts short action chunks instead of single steps. It learns from teleoperated data and often achieves high success rates.
This policy has been trained and pushed to the Hub using LeRobot.
See the full documentation at LeRobot Docs.
How to Get Started with the Model
For a complete walkthrough, see the training guide.
Below is the short version on how to train and run… See the full description on the dataset page: https://huggingface.co/datasets/mahmud8248/SO101_test109.so101-ball-cup-intervene-edited_6_v21so101-ball-cup-eval-intervene_7so101-ball-cup-intervene-edited_5so101-ball-cup-intervene-edited_4so101-ball-cup-intervene-edited_7_v21so101-ball-cup-eval-edited_6so101-press-button-base_v21so101_pp_donuts_v1_reward_videosso101-press-button-baseso101-candy-pickup
so101-candy-pickup
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
so101-dataset
SO-101 teleop cropped 224x224 dataset
This dataset mirrors the local SO-101 teleop HDF5 tree. Every HDF5 image
stream is stored under pixels/<view> with semantic camera names. Frames have
been cropped with saved crop metadata or the camera default crop and resized to
224x224 JPEG frames.
The original numeric datasets and timing/alignment metadata are preserved.
Semantic camera names are stored in semantic_camera_name attrs and summarized
in metadata/camera_names.json. The held-out… See the full description on the dataset page: https://huggingface.co/datasets/mo378/so101-dataset.so101-pick-placeso101_testso101-dataset-fullres
