datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
language_table_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "xarm",
"total_episodes": 442226,
"total_frames": 7045476,
"total_tasks": 127605,
"total_videos": 442226,
"total_chunks": 443,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:442226"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot.droid_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 92233,
"total_frames": 27044326,
"total_tasks": 31308,
"total_videos": 276699,
"total_chunks": 93,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:92233"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/droid_lerobot.FastUMI_100k_lerobot
FastUMI-100K: Advancing Data-Driven Robotic Manipulation with a Large-Scale UMI-Style Dataset
[paper] [dataset]
## Overview
FastUMI-100K is a large-scale, high-quality UMI-style dataset designed for data-driven robotic manipulation learning. Featuring over **100K+ demonstration trajectories** across **54 diverse tasks** and hundreds of object types, the dataset provides multi-view wrist-mounted fisheye images and high-frequency end-effector states. To… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/FastUMI_100k_lerobot.kuka_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "kuka_iiwa",
"total_episodes": 209880,
"total_frames": 2455879,
"total_tasks": 1,
"total_videos": 209880,
"total_chunks": 210,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:209880"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/kuka_lerobot.bridge_orig_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "widowx",
"total_episodes": 53192,
"total_frames": 1893026,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot.fractal20220817_data_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 87212,
"total_frames": 3786400,
"total_tasks": 599,
"total_videos": 87212,
"total_chunks": 88,
"chunks_size": 1000,
"fps": 3,
"splits": {
"train": "0:87212"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.robotwin-clean-and-aug-lerobot
Robotwin Dataset
Robotwin dataset in Lerobot format, with video latents already extracted in WAN 2.2 format, ready for use in Lingbot-VA post-training.
License Agreement
This project is licensed under the CC BY-NC-SA 4.0.
vlabench_primitive_pretrain_lerobot
Datacard
This is the official VLABench primitive pretraining dataset converted to the
LeRobot format. The dataset contains language-conditioned manipulation
trajectories collected with a Franka Panda robot in VLABench simulation.
This LeRobot version is hosted at:
https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot
Source
Project Page: https://vlabench.github.io/
Arxiv Paper: https://arxiv.org/abs/2412.18194
Code:… See the full description on the dataset page: https://huggingface.co/datasets/VLABench/vlabench_primitive_pretrain_lerobot.community_dataset_v3
Lerobot Community Datasets v3 - A Cross-Embodiment Pretraining Dataset for Vision Language Action Models
A large-scale robotics dataset for vision-language-action learning, featuring 791 datasets across 46 robot types, enabling cross-embodiment pretraining for generalist robot policies.
Overview
This is a crowdsourced, open-source dataset compiled from 235 community contributors worldwide. Building upon the pretraining datasets used for SmolVLA, Community… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/community_dataset_v3.EgoDex-LeRobot-v3.0EgoDex dataset from https://github.com/apple/ml-egodex, converted to LeRobot datasets v3.0 format.
This repo is structured with train / test directories at the top level,
under each of which are multiple subdirectories, each corresponding to one directory in the original EgoDex data format.
Each of these represents one LeRobot dataset, so to load the full EgoDex training set you have to load each subdirectory
as its own LeRobotDataset and then combine them (e.g. w/ PyTorch's ConcatDataset), or… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/EgoDex-LeRobot-v3.0.rbo_oxe_base_language_table_lerobot
Language Table (LeRobot) — Task-Pruned, Reindexed Subset
This release is a task-pruned subset of the original
IPEC-COMMUNITY/language_table_lerobot.
We subsampled by task text and rebuilt the package so it remains internally consistent
(indices, splits, stats, paths).
Robot: xArm
Modality: RGB video + states + actions
FPS / Resolution: 10 FPS, 360×640, AV1
License: apache-2.0 (inherits from source)
What’s different in this subset
Kept ~0.85% of unique tasks… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/rbo_oxe_base_language_table_lerobot.libero-assetsdroid_1.0.1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Franka",
"total_episodes": 95658,
"total_frames": 27630375,
"total_tasks": 49630,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 15,
"splits": {
"train": "0:95658"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/droid_1.0.1.libero_plus_lerobotaloha_sim_transfer_cube_humanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 20000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_transfer_cube_human.pushtThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 206,
"total_frames": 25650,
"total_tasks":1,
"total_videos": 206,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:206"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/pusht.bc_z_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 39350,
"total_frames": 5471693,
"total_tasks": 104,
"total_videos": 39350,
"total_chunks": 40,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39350"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bc_z_lerobot.HIW-500-LeRobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.images.head": {
"dtype": "video",
"shape": [
480,
1280,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height":… See the full description on the dataset page: https://huggingface.co/datasets/BitRobot/HIW-500-LeRobot.liberoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1693,
"total_frames": 273465,
"total_tasks": 40,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10.0,
"splits": {
"train": "0:1693"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/libero.bridge_v2_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "widowx",
"total_episodes": 53192,
"total_frames": 1999410,
"total_tasks": 19974,
"total_videos": 212768,
"total_chunks": 54,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:53192"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jesbu1/bridge_v2_lerobot.droid_rawvlabench-assetsInternData-A1-LeRobot-v3.0-by-embodimentInternData-A1 dataset taken from InternRobotics/InternData-A1,
with the tarballs extracted and directory structure "transposed" so that the top-level subdirectories are the four embodiments.
Two franka dirs
For the franka embodiment, there are two different feature spaces, so we split it into the franka-1 and franka-2 directories.
The feature spaces differ in image shape and gripper value range.
Minor fixes
Some subsets such as… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/InternData-A1-LeRobot-v3.0-by-embodiment.submissionsrobotwin_unifiedrecam-lerobotGalaxea-Open-World-Dataset-LeRobot-v3.0Galaxea Open-World Dataset taken from OpenGalaxea/Galaxea-Open-World-Dataset,
converted to LeRobot Datasets v3.0 format using lerobot.datasets.v30.convert_dataset_v21_to_v30.
Missing subsets
The subset Boil_The_Water_20250714_006 is missing due to the original files having some episodes at 62 fps,
which causes the conversion script to crash with an error.
The subset Put_The_Items_Into_The_Storage_Box_20250929_002_007 is missing due to it having 7 DoF arms rather than 6 DoF.… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/Galaxea-Open-World-Dataset-LeRobot-v3.0.aloha_sim_insertion_humanThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "aloha",
"total_episodes": 50,
"total_frames": 25000,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/aloha_sim_insertion_human.abc_130k_v3_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"left_arm_joint_1",
"left_arm_joint_2",
"left_arm_joint_3",
"left_arm_joint_4",
"left_arm_joint_5"… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/abc_130k_v3_train.VR-egodex-apple-lerobotEgoDex dataset from https://github.com/apple/ml-egodex, converted to LeRobot datasets v3.0 format.
This repo is structured with train / test directories at the top level,
under each of which are multiple subdirectories, each corresponding to one directory in the original EgoDex data format.
Each of these represents one LeRobot dataset, so to load the full EgoDex training set you have to load each subdirectory
as its own LeRobotDataset and then combine them (e.g. w/ PyTorch's ConcatDataset), or… See the full description on the dataset page: https://huggingface.co/datasets/bi199797/VR-egodex-apple-lerobot.
