datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
behavior1k-only-rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/k1000dai/behavior1k-only-rgb.wb_pointed_chair_pull_push_rgb
wb_pointed_chair_pull_push_rgb
Whole-body teleoperation data from a Unitree_G1_WholeBody_RGB, published in LeRobot v2.1 format.
Published in the v2.1 layout (one parquet and one video clip per episode) so it loads directly on older lerobot releases. On lerobot v3.0+ run the official upgrade first:
python -m lerobot.datasets.v30.convert_dataset_v21_to_v30 --repo-id=DaoyuanZhu/wb_pointed_chair_pull_push_rgb
Task — pull out the chair indicated by the human gesture, then push it… See the full description on the dataset page: https://huggingface.co/datasets/DaoyuanZhu/wb_pointed_chair_pull_push_rgb.demogen-rgbd-fr3-cube-stacking
demogen-rgbd — FR3 cube stacking, 10,000 augmented episodes
Augmented teleoperation data for a Franka FR3 cube-stacking task, produced by
demogen-rgbd: a DemoGen-style pipeline that turns one teleoperated episode
into many by translating the manipulated cube in 3-D and re-rendering the scene
from the original RGB-D observations.
No generative model is involved. Every pixel is a real observation, warped
by a dense inverse RGB-D transform. The only synthesised regions are the… See the full description on the dataset page: https://huggingface.co/datasets/Jordano/demogen-rgbd-fr3-cube-stacking.egocentric-stereo-rgbd
Hub Egocentric: Stereo RGB-D
9 egocentric stereo RGB-D clips with dense metric depth, IMU, 6-DoF VIO pose, and calibration. Each clip carries a self-contained LeRobot v3.0 dataset (loads on lerobot >= 0.6.0) plus side-by-side stereo, mono, a colorized depth preview, and a Foxglove MCAP recording.
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on StereoLabs ZED X Mini. Egocentric, human-demonstration data (passive; no robot action stream). July… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-stereo-rgbd.egocentric-gopro-rgb-imu
Hub Egocentric: GoPro RGB+IMU
19 egocentric human-manipulation clips captured on GoPro HERO13, with high-rate IMU (~200 Hz GPMF) delivered as CSV/Parquet/JSON sidecars plus a Foxglove MCAP recording.
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on GoPro HERO13. Egocentric, human-demonstration data (passive; no robot action stream). July 2026.
Dataset structure
Each clip is a top-level folder named #NN_... holding its media, a… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-gopro-rgb-imu.egocentric-iphone-rgb-imu
Hub Egocentric: iPhone RGB+IMU
12 egocentric human-manipulation clips captured on iPhone, with nominal 30 Hz CoreMotion + ARKit IMU/attitude on the shared media timeline, delivered as CSV/Parquet/JSON sidecars plus a Foxglove MCAP recording and per-clip camera intrinsics.
Across this 12-clip sample, IMU row count is 0–4 boundary rows lower than decoded video frame count (≤0.035%).
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on iPhone 13… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-iphone-rgb-imu.task1_1_5_rgb_recover_up_down_trim_194epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task1_1_5_rgb_recover_up_down_trim_194ep.robotwin_blocks_ranking_rgb_hybrid_200_dynFcam
robotwin_blocks_ranking_rgb_hybrid_200_dynFcam
A validated LeRobot v2.1 release of two native RoboTwin 2.0 expert schedules for blocks_ranking_rgb.
The variants use independent accepted seeds and are concatenated into one training split; matching episode offsets are not paired scenes.
dynFcam observation derivative
This repository reuses exactly the same validated native episodes as Shiki42/robotwin_blocks_ranking_rgb_hybrid_200 at revision… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/robotwin_blocks_ranking_rgb_hybrid_200_dynFcam.table_spill_cleanup_bimanual_rgbd_segmentation_poses
Exylos Bimanual Table Spill Cleanup Rich-Modality Sample
A compact, rich-modality bimanual robot manipulation dataset for tabletop spill cleanup.
Each episode combines synchronized dual-arm Panda state/action trajectories, 7 RGB camera streams, per-frame depth maps, per-frame segmentation masks, object pose streams, phase annotations, and an objective cleanup success metric based on the remaining spill fraction.
This dataset is a rich-modality inspection sample for the Exylos… See the full description on the dataset page: https://huggingface.co/datasets/ExylosAi/table_spill_cleanup_bimanual_rgbd_segmentation_poses.every_frame_right_rgbdice_new_rgb_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 6098,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jy13/dice_new_rgb_test.task1_1_5_task1_1_6_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task1_1_5_task1_1_6_rgb.task2_1_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task2_1_rgb.task4_2_2_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task4_2_2_rgb.task4_3_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task4_3_rgb.task1_1_6_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task1_1_6_rgb.task4_1_task4_2_2_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task4_1_task4_2_2_rgb.keys_into_bowl_ob15_depth_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"names": [
"left_arm_shoulder_pan.pos",
"left_arm_shoulder_lift.pos",
"left_arm_elbow_flex.pos",
"left_arm_wrist_flex.pos",
"left_arm_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/Odog16/keys_into_bowl_ob15_depth_rgb.game-scenes-posed-rgbd
Origin Lab Game Scenes: Posed RGB-D Flythroughs of Game Worlds
Every frame carries the camera that rendered it and the depth the engine computed for it. Ten game worlds, with the camera released from the player for 60% of the footage: metric depth, world-space normals, 4x4 pose, and per-frame intrinsics on one frame index, plus hundreds of full in-place turns and long stretches in which the world is frozen and only the camera moves. Two trajectories per world and one whole… See the full description on the dataset page: https://huggingface.co/datasets/originlab/game-scenes-posed-rgbd.train-rgb-all-depth-eyeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 1,
"total_frames": 245,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hngchris/train-rgb-all-depth-eye.maniskill3-sft-rgb-lerobot-1200epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "maniskill3_sim",
"total_episodes": 1200,
"total_frames": 171480,
"total_tasks": 6,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null,
"features":… See the full description on the dataset page: https://huggingface.co/datasets/onnoboru/maniskill3-sft-rgb-lerobot-1200ep.task4_2_2_task4_3_rgbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task4_2_2_task4_3_rgb.droid-3d-rgb-68epThis is a FiftyOne dataset with 68 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/droid-3d-rgb-68ep")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for droid_3d (68-episode FiftyOne RGB subset)
A… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/droid-3d-rgb-68ep.multispecqr-rgb-datasettask1_1_5_rgb_recover_up_down_trim_174epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task1_1_5_rgb_recover_up_down_trim_174ep.task1_1_5_rgb_recover_up_down_trim_204epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task1_1_5_rgb_recover_up_down_trim_204ep.pick-up-erasers-place-in-drawers-rgbd-v2-fixed-mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_yaw.pos",
"wrist_roll.pos",
"gripper.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/namin72/pick-up-erasers-place-in-drawers-rgbd-v2-fixed-merged.task1_1_5_rgb_recover_up_downThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"arm_left_joint_0",
"arm_left_joint_1",
"arm_left_joint_2",
"arm_left_joint_3",
"arm_left_joint_4"… See the full description on the dataset page: https://huggingface.co/datasets/takeru01/task1_1_5_rgb_recover_up_down.soarm_pp_rgb_random_pose_red_target_100epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 37665,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lanyard2400/soarm_pp_rgb_random_pose_red_target_100ep.SO101-ma_stack_RGBblock_on_bluedish_30fpsThis dataset was created using LeRobot.
Dataset Description
SO-101 robot dataset collected with the MA autonomous forward/reset pipeline. The task is to stack red, green, and blue RGB blocks on the blue dish from bottom to top. The dataset contains 100 successful episodes at 30 FPS with top and left-wrist RGB camera streams, robot state/action features, end-effector pose, gripper state, and skill/subtask annotations.
Homepage: https://huggingface.co/CoRL2026-CSI
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/SO101-ma_stack_RGBblock_on_bluedish_30fps.
