datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-Robotics-GR00T-Teleop-GR1
Introduction
TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments.
Project page: https://dreamdojo-world.github.io/
Paper: https://arxiv.org/abs/2602.06949
Code: https://github.com/NVIDIA/DreamDojo
How to Use
Check out https://github.com/NVIDIA/DreamDojo
Citation
@article{gao2026dreamdojo,
title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.GR1-Tabletop-NextState-1000x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-NextState-1000x24.gr1_mg_gr00t_300_new
LeRobot dataset: gr1_mg_gr00t_300_new
Folder layout:
data/ – tensors, arrays, per-episode artifacts
meta/ – episodes.jsonl, tasks.jsonl, stats.json, modality.json, info.json
videos/ – episode videos (if any)
This is a raw LeRobot-format dataset (not a 🤗 Datasets script). Download via git lfs or huggingface_hub.snapshot_download.
gr1_arms_waist-CuttingboardToPangr1_arms_waist-CuttingboardToCardboardBoxgr1_arms_waist-PlaceMilkToMicrowavegr1_arms_waist-WineToCabinetgr1_arms_waist-TrayToTieredShelfgr1_arms_waist-TrayToPlategr1_arms_waist-TrayToPotgr1_arms_waist-PlateToCardboardBoxgr1_arms_waist-PlateToPangr1_arms_waist-CupToDrawergr1_arms_waist-PlaceBottleToCabinetgr1_arms_waist-PlateToBowlgr1_arms_waist-PotatoToMicrowave_gr1_unified
PhysicalAI-Robotics-GR00T-X-Embodiment-Sim
Github Repo: Isaac GR00T N1
We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks.
Cross-embodied bimanual manipulation: 9k trajectories
Dataset Name
#trajectories
bimanual_panda_gripper.Threading
1000
bimanual_panda_hand.LiftTray
1000
bimanual_panda_gripper.ThreePieceAssembly
1000… See the full description on the dataset page: https://huggingface.co/datasets/nhatchung/_gr1_unified.gr1_arms_waist-PlacematToTieredShelfArena-GR1-Manipulation-PlaceItemCloseDoor-Task
Dataset Description:
The Arena-GR1-Manipulation-PlaceItemCloseDoor-Task dataset is a multimodal collection of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation tasks in the IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, and action) needed to train and evaluate generalist robot policies for a sequential task (e.g. putting object into a fridge and closing the door).
Dataset Name
# Trajectories
GR1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-PlaceItemCloseDoor-Task.GR1-Tabletop-Merged-100x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-Merged-100x24.Arena-GR1-Manipulation-Task
Dataset Description:
The Arena-GR1-Manipulation-Task dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) manipulation task in IsaacLab-Arena environment. Each entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for opening microwave task.
Dataset Name
# Trajectories
GR1 Manipulation Task
50
This dataset is ideal for behavior cloning, policy learning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Arena-GR1-Manipulation-Task.GR1-Tabletop-NextState-100x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-NextState-100x24.gr1_arms_waist-PlateToPlateGR1-Tuned-Tasks
Dataset Description:
This dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) tabletop manipulation tasks for industrial settings. Each dataset entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for tasks like pouring nuts or sorting pipes by color.
Dataset Name
# Trajectories
Exhaust-Pipe-Sorting-task
1000
Nut-Pouring-task
1000
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/benlam62/GR1-Tuned-Tasks.robocasa_gr1_tabletop_tasks_fingertipsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "robomimic",
"total_episodes": 2235,
"total_frames": 552031,
"total_tasks": 20,
"total_videos": 2235,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:2235"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ishika/robocasa_gr1_tabletop_tasks_fingertips.robocasa_gr1_tabletop_tasksThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "robomimic",
"total_episodes": 2235,
"total_frames": 552031,
"total_tasks": 20,
"total_videos": 2235,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1877",
"test": "1877:2235"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ishika/robocasa_gr1_tabletop_tasks.gr1_arms_waist-CuttingboardToTieredBasketgr1_arena_sequential_task_replay
GR1 Arena — Ranch Bottle Into Fridge (new camera pose, ego + wrist, replay)
LeRobot-format teleoperation/replay dataset for the GR1 humanoid performing the
put_item_in_fridge_and_close_door task in Isaac Lab Arena.
Task: Place the ranch dressing bottle on the top shelf of the fridge, and
close the fridge door. (object: ranch_dressing_hope_robolab)
What this dataset is
This is a re-rendered / replayed version of the official NVIDIA Arena
dataset. The source… See the full description on the dataset page: https://huggingface.co/datasets/china-sae-robotics/gr1_arena_sequential_task_replay.GR1-Tabletop-Merged-300x24
GR1 Tabletop Merged LeRobot Datasets
Merged and subsampled versions of the GR1 tabletop manipulation datasets from the NVIDIA PhysicalAI-Robotics-GR00T-X-Embodiment-Sim collection, formatted in LeRobot v2.0 format.
Dataset Variants
Variant
Demos/Task
Tasks
Total Episodes
Total Frames
Approx Size
1000x24/
1000
24 folders, 186 unique tasks
24,000
6,020,058
~40 GB
300x24/
300
24 folders, 186 unique tasks
7,200
1,803,236
~12 GB
100x24/
100
24 folders… See the full description on the dataset page: https://huggingface.co/datasets/Joocjun/GR1-Tabletop-Merged-300x24.tavis-head-gr1t2-800epThis dataset was created using LeRobot and
is presented here as a FiftyOne dataset.
Installation
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
dataset = load_from_hub("Voxel51/tavis-head-gr1t2-800ep")
session = fo.launch_app(dataset)
Dataset Card for TAVIS Head GR1T2 (800-episode FiftyOne export)
Dataset Details
Dataset Description
TAVIS Head GR1T2 is part… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/tavis-head-gr1t2-800ep.
