datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-Robotics-Locomanipulation-GRAIL
📢 News
[2026-07-15] Released task-general tracking policy checkpoints trained on the released data. Follow the tracking doc to use them to track our released motion data.
[2026-07-14] Updated data/pickup_table and data/pickup_ground. If you downloaded them before this date, please re-download.
Dataset Overview
Tabletop Pickup
Ground Pickup
Tabletop Manipulation
Ground Manipulation
Sitting
Curb
Slope… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Locomanipulation-GRAIL.PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.PhysicalAI-Robotics-NuRec
Dataset Description
The Physical AI NuRec dataset seeks to empower robotic researchers to build the next generation of physical AI based end-to-end robotic models.
This dataset includes various 3DGUT in USD files that can be loaded in Isaac Sim. Some datasets also include a mesh and occupancy map. The Mesh components are used for collision detection while the 3DGUT components provide realistic rendering. The asset can also be used with Isaac Sim Extensions like MobilityGen for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-NuRec.Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.PhysicalAI-VANTAGE-Bench
VANTAGE-BENCH
Video ANalysis Tasks Across Generalized Environments
Paper: VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Dataset Description
VANTAGE-BENCH is the first public benchmark purpose-built for evaluating visual understanding on video captured by fixed infrastructure cameras. It spans three real-world domains — warehouse, smart city / Intelligent Transportation Systems (ITS), and smart spaces — across six spatio-temporal… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-VANTAGE-Bench.forge-arm-pixels
physicalai-bmi/forge-arm-pixels
Real MuJoCo pixels captured live from the Institute's in-browser Forge arm (WebGPU),
paired with the action the released state-checkpoint took. This is the exact training
set behind physicalai-bmi/nano-vla-pixels.
2,500 frames across 128 reaches, frames/f#####.png (the rendered MuJoCo arm, 844×520).
meta.json — per-frame { i, act:[3], obs:[7], reaches }; act is the 3-D joint-delta action, reaches is the episode index (use it for an episode-level… See the full description on the dataset page: https://huggingface.co/datasets/physicalai-bmi/forge-arm-pixels.isaac-sdg-rescue-target
isaac-sdg-rescue-target
Synthetic training data for detecting a person lying on the ground wearing a hi-vis vest,
the target class of the search-and-rescue spotter. Rendered with NVIDIA Isaac Sim 6.0.1
Replicator (path-traced RTX) from posed, vested worker characters placed in real Isaac
environments, with per-frame randomization of camera position, HDRI sky, key light,
target pose, and distractor objects with random materials. Ground truth is exact: boxes come
from the renderer… See the full description on the dataset page: https://huggingface.co/datasets/ubr-physical-ai/isaac-sdg-rescue-target.PhysicalAI-Robotics-GR00T-Eval
EVAL-175
Dataset Description:
123 initial frame pictures from the robot's perspective before performing various tasks in the lab.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA Corporation (GEAR Lab)
Dataset Creation Date:
May 1, 2025
License/Terms of Use:
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
This dataset was created using a… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Eval.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.ubr-maze-nav
ubr-maze-nav — vision + instrument waypoint planning for a small tracked robot
Synthetic navigation corpus for fine-tuning small vision-language models to
plan local waypoint paths for a 0.3 m tracked ground robot in corridor/maze
environments, plus the frozen evaluation suite used in our internal reports.
Each sample is one first-person RGB frame (640×480) from the robot's camera
in a procedurally generated MuJoCo scene, an instruction carrying the goal
(bearing/range) and a… See the full description on the dataset page: https://huggingface.co/datasets/ubr-physical-ai/ubr-maze-nav.physical-ai-bench-understanding-evalPhysicalAI-AV-Counterfactual
PhysicalAI-AV-Counterfactual
A curated evaluation dataset for autonomous vehicle (AV) safety research. Each sample pairs a real forward-facing dashcam frame with a generative-AI-edited counterfactual in which a plausible driving hazard has been inserted, along with the ego vehicle's past and future trajectory extracted from the source recording.
The dataset is designed to benchmark whether a vision-language model (VLM) can correctly identify and reason about a newly-appeared… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/PhysicalAI-AV-Counterfactual.nvidia-physical-ai-keyframes-sample
Dataset Card for 2025.11.20.21.48.24.136991
This is a FiftyOne dataset with 1000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("dgural/nvidia-physical-ai-keyframes-sample")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/dgural/nvidia-physical-ai-keyframes-sample.PhysicalAI-Robotics-mindmap-GR1-Drill-in-Box
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Drill in Box task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-GR1-Drill-in-Box.physicalai-simready-warehouse-01-test
NVIDIA Physical AI Warehouse OpenUSD Dataset
Dataset Version: 0.1.0
Date: March 18, 2025
Author: NVIDIA, Corporation
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_Warehouse_01.csv). The CSV file is organized in the following manner:
asset_name : Name of the asset
relative_path: location… See the full description on the dataset page: https://huggingface.co/datasets/nvaszabo/physicalai-simready-warehouse-01-test.PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Cube Stacking task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking.PhysicalAI-Robotics-mindmap-Franka-Mug-in-Drawer
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Mug in Drawer task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-Franka-Mug-in-Drawer.PhysicalAI-SimReady-Homes
PhysicalAI SimReady Homes: Multi-Room Interiors
Version 1.0 · 1,000 simulation-ready multi-room home interiors that load into Isaac
Sim and are ready to train on — no rigging, no retopology, no physics authoring.
Every scene is a furnished multi-room home in OpenUSD with rigid bodies, mass and inertia,
collision approximations, PhysX friction/restitution/density, articulated doors and drawers,
per-object semantic labels, PBR materials, HDRI lighting and placed cameras — plus a… See the full description on the dataset page: https://huggingface.co/datasets/imagineio/PhysicalAI-SimReady-Homes.physicalAIPhysicalAI-US-Evaluation
PhysicalAI-US-Evaluation
A held-out US evaluation set for the navigation planner: 19,744 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn — uniformly at random — from the pool of US scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-US-Evaluation.Physical-AI-AV-FR
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,909 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-FR.PhysicalAI-DE-Evaluation
PhysicalAI-DE-Evaluation
A held-out German evaluation set for the navigation planner: 19,999 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn from the pool of German scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT) stage, and
the… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-DE-Evaluation.Physical-AI-AV-ES
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,674 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-ES.PhysicalAI-AV-Counterfactual
Dataset Structure
Each scene is stored under its SceneID directory. Within each scene directory, samples are grouped by timestamp. Each timestamp may have up to three associated files:
<SceneID>/
<timestamp>.png
<timestamp>.pkl
<timestamp>-nano-banana.png
File Descriptions
File
Description
<timestamp>.png
Forward-facing camera frame nearest to the requested timestamp.
<timestamp>.pkl
Pickle file containing scene metadata and ego-motion trajectory… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-AV-Counterfactual.Physical-AI-AV-DE
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 324,105 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 10 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.Physical-AI-AV-IT
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,991 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.
