datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is organized in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-SimReady-Warehouse-01.PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.PhysicalAI-VANTAGE-Bench
VANTAGE-BENCH
Video ANalysis Tasks Across Generalized Environments
Paper: VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Dataset Description
VANTAGE-BENCH is the first public benchmark purpose-built for evaluating visual understanding on video captured by fixed infrastructure cameras. It spans three real-world domains — warehouse, smart city / Intelligent Transportation Systems (ITS), and smart spaces — across six spatio-temporal… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-VANTAGE-Bench.physical-ai-bench-conditional-generation
Physical AI Bench - Conditional Generation
Paper | Code
This dataset (Phsical AI benchmark, PAI-Bench) consisting of 600 examples across three key scenarios: robotic arm operations, driving, and ego-centric everyday life scenes, each representing a critical aspect of Physical AI. This dataset is constructed by sampling a number of videos from three different datasets. The specific details are provided below.
Dataset
Category
Sample Nums
Agibot World
Robotics
200
OpenDV… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-conditional-generation.PhysicalAI-Robotics-GraspGen
GraspGen: Scaling Sim2Real Grasping
GraspGen is a large-scale simulated grasp dataset for multiple robot embodiments and grippers.
We release over 57 million grasps, computed for a subset of 8515 objects from the Objaverse XL (LVIS) dataset. These grasps are specific to three grippers: Franka Panda, the Robotiq-2f-140 industrial gripper, and a single-contact suction gripper (30mm radius).
Dataset Format
The dataset is released in the WebDataset format. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GraspGen.physical-ai-bench-understanding
Physical AI Bench - Understanding
PAI-Bench (Physical AI Bench) is a comprehensive benchmark designed to evaluate physical AI generation and understanding capabilities across various real-world scenarios. This particular dataset, PAI-Bench-U, focuses specifically on Video Understanding tasks, comprising 2,808 real-world cases with task-aligned metrics.
Paper: PAI-Bench: A Comprehensive Benchmark For Physical AI
Code: GitHub Repository
Citation
If you use Physical AI… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-understanding.Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval
VoMP: Predicting Volumetric Mechanical Properties
Dataset Description:
The Pre-Processed 3D Dataset is a dataset that is composed of 4 individual 3D asset datasets which are processed to render them from multiple views, voxelize the assets, and propagate VLM annotations for material properties.
We release pre-processed data derived from the 3D assets, specifically: voxels, rendered images, and LLM-annotated material descriptions.
This dataset is for research and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-PhysicalAssets-VoMP-Eval.PhysicalAI-Robotics-GR00T-Eval
EVAL-175
Dataset Description:
123 initial frame pictures from the robot's perspective before performing various tasks in the lab.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA Corporation (GEAR Lab)
Dataset Creation Date:
May 1, 2025
License/Terms of Use:
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
This dataset was created using a… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Eval.PhysicalAI-Robotics-GR00T-GR1
GR1-100
92 Videos of a Fourier GR1-T2 robot performing various tasks in the lab from the third person perspective.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA Corporation (GEAR Lab)
Dataset Creation Date:
May 1, 2025
License/Terms of Use:
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
This dataset was created using a Fourier GR1-T2 robot.
Intended… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-GR1.PhysicalAI-SimReady-Warehouse-01
NVIDIA Physical AI SimReady Warehouse OpenUSD Dataset
Dataset Version: 1.1.0
Date: May 18, 2025
Author: NVIDIA, Corporation
License: CC-BY-4.0 (Creative Commons Attribution 4.0 International)
Contents
This dataset includes the following:
This README file
A CSV catalog that enumerates all of the OpenUSD assets that are part of this dataset including a sub-folder of images that showcase each 3D asset (physical_ai_simready_warehouse_01.csv). The CSV file is… See the full description on the dataset page: https://huggingface.co/datasets/yzhllm/PhysicalAI-SimReady-Warehouse-01.punch-energetics
Punch Energetics
Everything needed to check, rerun, or disagree with TR-2026-41, Joules per Punch
(Institute for Physical AI @ John Bailey Institute, The Charlot Lab).
Eleven single-file Rust benches, no dependencies, and the thirteen run receipts every published
figure in TR-2026-41 and TR-2026-45 is taken from.
The question
A humanoid throws a cross. Roughly three-quarters of a human straight punch comes from leg
drive and trunk rotation, and a Unitree G1 has… See the full description on the dataset page: https://huggingface.co/datasets/physicalai-bmi/punch-energetics.PhysicalAI-ADE-US
PhysicalAI-US-ADE
Dataset Summary
PhysicalAI-US-ADE contains per-sample evaluation outputs for autonomous driving waypoint prediction on the US subset of the PhysicalAI NVIDIA dataset.
This dataset stores inference-time predictions and evaluation statistics for models evaluated on the dataset, organized by model name at the top level. Each model directory contains sample-level records for that model’s predictions against ground truth.
The current release includes… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-ADE-US.physical-ai-bench-understanding-evalphysical-ai-bench-artifactsPhysicalAI-Robotics-mindmap-GR1-Drill-in-Box
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Drill in Box task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-GR1-Drill-in-Box.PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Cube Stacking task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking.PhysicalAI-ADE-DEPhysicalAI-Robotics-mindmap-Franka-Mug-in-Drawer
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Mug in Drawer task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-Franka-Mug-in-Drawer.PhysicalAI-SimReady-Homes
PhysicalAI SimReady Homes: Multi-Room Interiors
Version 1.0 · 1,000 simulation-ready multi-room home interiors that load into Isaac
Sim and are ready to train on — no rigging, no retopology, no physics authoring.
Every scene is a furnished multi-room home in OpenUSD with rigid bodies, mass and inertia,
collision approximations, PhysX friction/restitution/density, articulated doors and drawers,
per-object semantic labels, PBR materials, HDRI lighting and placed cameras — plus a… See the full description on the dataset page: https://huggingface.co/datasets/imagineio/PhysicalAI-SimReady-Homes.physical-ai-evals-libero-spatial-pilot
LIBERO-Spatial paired pilot
Historical rollout records for OpenVLA and VLA-JEPA on 100 paired LIBERO-Spatial episode
specifications. The dataset contains 200 policy/episode records and 23,283 transition rows.
This is an exploratory trace, not a LIBERO reference evaluation or a confirmatory model
comparison. The artifact label 2026-07-02 is not a recorded execution timestamp.
Files
steps.parquet: normalized rollout-v1 transition rows.
episodes.parquet: one row per… See the full description on the dataset page: https://huggingface.co/datasets/Eventual-Inc/physical-ai-evals-libero-spatial-pilot.PhysicalAI-US-Evaluation
PhysicalAI-US-Evaluation
A held-out US evaluation set for the navigation planner: 19,744 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn — uniformly at random — from the pool of US scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-US-Evaluation.Physical_AIphysicalai-ood-submissions-pending
Pending submissions for NVIDIA PhysicalAI-OOD-Leaderboard
HF free-tier cannot call the private ZeroGPU evaluator (duration=450s > free max).
These JSON packs are format-validated against /submit (fail only at GPU step).
Request for @WenyanCong / organizers
Please run evaluation for HF user VeigaPunk using these files (or lower evaluator duration so free users can submit):
File
Track
Notes
submission_track2_val_gold.json
track2
349 keys, public CoC style… See the full description on the dataset page: https://huggingface.co/datasets/VeigaPunk/physicalai-ood-submissions-pending.Physical-AI-AV-FR
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,909 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-FR.PhysicalAI-DE-Evaluation
PhysicalAI-DE-Evaluation
A held-out German evaluation set for the navigation planner: 19,999 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn from the pool of German scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT) stage, and
the… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-DE-Evaluation.PhysicalAI-SimReady-Kitchens-v1
PhysicalAI SimReady Kitchens: 800 Scenes
800 simulation-ready OpenUSD kitchen environments for robotics, embodied AI, and physical AI evaluation.
This release is a public sample of Imagine.io's programmable world-generation infrastructure. It is designed to help research and commercial teams evaluate controlled, metadata-rich indoor environments for perception, scene understanding, simulation, and physical AI workflows.
License notice. Public access is CC BY-NC 4.0 for… See the full description on the dataset page: https://huggingface.co/datasets/imagineio/PhysicalAI-SimReady-Kitchens-v1.Physical-AI-AV-ES
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,674 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-ES.Physical-AI-AV-DE
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 324,105 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 10 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.
