datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SmartSpaces
Physical AI Smart Spaces Dataset
Overview
Comprehensive, annotated dataset for multi-camera tracking and 2D/3D object detection. This dataset is synthetically generated with Omniverse and Cosmos Transfer.
This dataset consists of over 280 hours of video from across nearly 1,800 cameras from indoor scenes in warehouses, hospitals, retail, and more. The dataset is time synchronized for tracking humans, forklifts, pallet trucks and Autonomous Mobile Robots (AMRs)… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SmartSpaces.PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
Dataset Description:
PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.PhysicalAI-Robotics-Locomanipulation-GRAIL
📢 News
[2026-07-15] Released task-general tracking policy checkpoints trained on the released data. Follow the tracking doc to use them to track our released motion data.
[2026-07-14] Updated data/pickup_table and data/pickup_ground. If you downloaded them before this date, please re-download.
Dataset Overview
Tabletop Pickup
Ground Pickup
Tabletop Manipulation
Ground Manipulation
Sitting
Curb
Slope… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL.PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
PhysicalAI SDG-Warehouse
PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
Dataset Description:
The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research.
Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.PhysicalAI-Robotics-GR00T-Teleop-Sim
Simulation GR1 Tabletop Task 1K Dataset
Dataset Description:
The PhysicalAI-Robotics-GR00T-Teleop-GR1 dataset consists of 1000 teleoperation trajectories in simulation using the GR1 robot with upper body control. The simulation setup mimics tabletop manipulation tasks and uses RGB observations with a virtual camera. The robot is equipped with simulated Fourier hands.
This dataset is ready for non-commercial use.
Dataset Owner(s):
NVIDIA GEAR… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim.PhysicalAI-Robotics-GR00T-Teleop-GR1
Introduction
TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments.
Project page: https://dreamdojo-world.github.io/
Paper: https://arxiv.org/abs/2602.06949
Code: https://github.com/NVIDIA/DreamDojo
How to Use
Check out https://github.com/NVIDIA/DreamDojo
Citation
@article{gao2026dreamdojo,
title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card
Dataset Description
PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.PhysicalAI-VANTAGE-Bench
VANTAGE-BENCH
Video ANalysis Tasks Across Generalized Environments
Paper: VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Dataset Description
VANTAGE-BENCH is the first public benchmark purpose-built for evaluating visual understanding on video captured by fixed infrastructure cameras. It spans three real-world domains — warehouse, smart city / Intelligent Transportation Systems (ITS), and smart spaces — across six spatio-temporal… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-VANTAGE-Bench.physical-ai-bench-step-30000-videos
PhysicalAIBench Step 30000 Model Comparison
This repository contains two filename-aligned sets of 5,220 MP4 outputs from
step_30000 evaluation runs. The files are presented through the
companion PhysicalAI Video Gallery.
Model sets
Directory
Model
Files
Bytes
videos/
DC-AE v0.2, Cosmos encoder + causal decoder
5,220
4,244,965,541
videos_wan22_vae/
Wan 2.2 VAE, phase 3 PDX
5,220
4,818,900,749
The two directories have an exact 1:1 basename match.… See the full description on the dataset page: https://huggingface.co/datasets/spongy/physical-ai-bench-step-30000-videos.physical-ai-bench-conditional-generation
Physical AI Bench - Conditional Generation
Paper | Code
This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below.
Dataset
Category
Sample Nums
Agibot World
Robotics
200
OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.PhysicalAI-Robotics-Manipulation-Kitchen
PhysicalAI Robotics Manipulation in the Kitchen
Dataset Description:
PhysicalAI-Robotics-Manipulation-Kitchen is a dataset of automatic generated motions of robots performing operations such as opening and closing cabinets, drawers, dishwashers and fridges. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Kitchen.PhysicalAI-Robotics-GR00T-Teleop-G1
Unitree G1 Fruits Pick and Place 1K Dataset
Dataset Description:
The PhysicalAI-Robotics-GR00T-Teleop-G1 dataset consists of1000 teleoperation trajectories of real robot data using Unitree G1, with upper body control. The robot chooses the correct fruit to pick and place on the plate according to the language prompt. A total of 4 fruits are used: Apple, Pear, Starfruit, Grape. The robot is equipped with the default realsense camera, and a pair of Unitree G1 Tri-fingers… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-G1.physical-ai-bench-understanding
Physical AI Bench - Understanding
PAI-Bench (Physical AI Bench) is a comprehensive benchmark designed to evaluate physical AI generation and understanding capabilities across various real-world scenarios. This particular dataset, PAI-Bench-U, focuses specifically on Video Understanding tasks, comprising 2,808 real-world cases with task-aligned metrics.
Paper: PAI-Bench: A Comprehensive Benchmark For Physical AI
Code: GitHub Repository
Citation
If you use Physical AI… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-understanding.PhysicalAI-Traffic-Anomaly-Reasoning
Traffic Anomaly Reasoning (TAR)
This dataset is the official release for AI City Challenge 2026
Track 3 — Anomalous Events in Transportation.
It contains:
Training (train/): 44,040 pseudo-labeled multi-task annotations
covering 3,670 CCTV videos (≈26.1 hours: 9.2 hrs anomalous + 16.9 hrs
normal) sourced from eight public datasets.
Test (test/): 960 human-curated annotations covering 80 short
clips trimmed from 17 public YouTube videos. Answers are redacted in
this release;… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Traffic-Anomaly-Reasoning.Physical_AI_SO101_Cup_Nesting_Task
DecisionFacts Physical AI Dataset — SO-101 Robotic Arm Teleoperation
Data Summary
This dataset is a curated collection of real-world teleoperation data captured on the SO-101 robotic arm (so_follower), built to support training and evaluation of modern robot-learning models — from imitation-learning policies to large-scale Vision-Language-Action (VLA) and world models.
Each episode is a human-teleoperated demonstration of a manipulation task, recorded… See the full description on the dataset page: https://huggingface.co/datasets/DecisionFacts/Physical_AI_SO101_Cup_Nesting_Task.PhysicalAI-Robotics-Manipulation-ObjectsPhysicalAI-Robotics-Manipulation-Objects is a dataset of automatic generated motions of robots performing operations such as picking and placing objects in a kitchen environment. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with Kinova Gen3 arms. The environments are kitchen scenes where the furniture and appliances were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Manipulation-Objects.PhysicalAI-GR00T-Tuned-Tasks
Dataset Description:
This dataset is multimodal collections of trajectories generated in Isaac Lab. It supports humanoid (GR1) tabletop manipulation tasks for industrial settings. Each dataset entry provides the full context (state, vision, language, action) needed to train and evaluate generalist robot policies for tasks like pouring nuts or sorting pipes by color.
Dataset Name
# Trajectories
Exhaust-Pipe-Sorting-task
1000
Nut-Pouring-task
1000
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-GR00T-Tuned-Tasks.PhysicalAI-Event-Videos
PhysicalAI-Event-Videos
PhysicalAI-Event-Videos is a video-anomaly and event dataset for developing text-to-video anomaly-search and safety-event-understanding systems. Version 1.0 provides structured annotations for 1,612 parent-video records and 3,796 labeled chunks, with 71,023 captions and queries across person-attribute, general-caption, anomaly, and action search tasks.
The redistributable media payload contains 1,486 NVIDIA-generated parent videos and 2,939 packaged chunk… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Event-Videos.PhysicalAI-Robotics-GR00T-GR1
GR1-100
92 Videos of a Fourier GR1-T2 robot performing various tasks in the lab from the third person perspective.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA Corporation (GEAR Lab)
Dataset Creation Date:
May 1, 2025
License/Terms of Use:
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
This dataset was created using a Fourier GR1-T2 robot.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-GR1.PhysicalAI-Autonomous-Vehicles-Sample
Dataset Card for nvidia-physical-ai-sample
This dataset is a small curated sample (100 items) extracted from the fullNVIDIA PhysicalAI Autonomous Vehicles dataset.It is intended for quick experimentation, tutorials, and FiftyOne integration demos without requiring the multi-terabyte original dataset.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import… See the full description on the dataset page: https://huggingface.co/datasets/dgural/PhysicalAI-Autonomous-Vehicles-Sample.physical-ai
DecisionFacts Physical AI Dataset — SO-101 Robotic Arm Teleoperation
Data Summary
This dataset is a curated collection of real-world teleoperation data captured on the SO-101 robotic arm (so_follower), built to support training and evaluation of modern robot-learning models — from imitation-learning policies to large-scale Vision-Language-Action (VLA) and world models.
Each episode is a human-teleoperated demonstration of a manipulation task, recorded… See the full description on the dataset page: https://huggingface.co/datasets/DecisionFacts/physical-ai.physicalaiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/slvrfivo/physicalai.Physical_AI_SO101_Cup_Nesting_Task
DecisionFacts Physical AI Dataset — SO-101 Robotic Arm Teleoperation
Data Summary
This dataset is a curated collection of real-world teleoperation data captured on the SO-101 robotic arm (so_follower), built to support training and evaluation of modern robot-learning models — from imitation-learning policies to large-scale Vision-Language-Action (VLA) and world models.
Each episode is a human-teleoperated demonstration of a manipulation task, recorded… See the full description on the dataset page: https://huggingface.co/datasets/salehin21/Physical_AI_SO101_Cup_Nesting_Task.so101_1200ep_dataset_20260803_104129This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/physicalairi/so101_1200ep_dataset_20260803_104129.physical_ai_tuneupPhysicalAI-Robotics-Manipulation-ObjectsPhysicalAI-Robotics-Manipulation-Objects is a dataset of automatic generated motions of robots performing operations such as picking and placing objects in a kitchen environment. The dataset was generated in IsaacSim leveraging reasoning algorithms and optimization-based motion planning to find solutions to the tasks automatically [1, 3]. The dataset includes a bimanual manipulator built with Kinova Gen3 arms. The environments are kitchen scenes where the furniture and appliances were… See the full description on the dataset page: https://huggingface.co/datasets/muhammadshihab/PhysicalAI-Robotics-Manipulation-Objects.physical-ai-bucharest-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 594,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vladfatu/physical-ai-bucharest-1.ur10e-gear-pick-place-demos
UR10e Gear Pick-and-Place Demonstrations
Private archive of the exact LeRobot v3.0 train and validation directories used by the revised Rho delta-action IL run. The task is to pick up the large gear and place it in the bin using a simulated UR10e with a Robotiq 2F-140 gripper.
📊 Dataset identity
Split
Episodes
Frames
Original data bytes
train/
90
154,593
277,712,776
validation/
10
17,170
30,883,658
Total
100
171,763
308,596,434
The corpus was… See the full description on the dataset page: https://huggingface.co/datasets/physical-ai-toolchain/ur10e-gear-pick-place-demos.ur3_stack_cube_camera_sim_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
12
],
"names": [
"shoulder_pan_joint.pos",
"shoulder_lift_joint.pos",
"elbow_joint.pos",
"wrist_1_joint.pos",
"wrist_2_joint.pos"… See the full description on the dataset page: https://huggingface.co/datasets/physicalairi/ur3_stack_cube_camera_sim_v2.
