datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.Physical-AI-AV-US
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 2,789,773 samples from 150 000
driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the
United States.
Format
WebDataset — 100 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.png
Front-facing wide-angle camera frame (640 × 360 px)
{key}.json
Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.physical-ai-bench-artifactsPhysicalAI-Robotics-mindmap-GR1-Drill-in-Box
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Drill in Box task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-GR1-Drill-in-Box.PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Cube Stacking task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-Franka-Cube-Stacking.PhysicalAI-Robotics-mindmap-Franka-Mug-in-Drawer
Dataset Description:
This dataset is a multimodal collection of trajectories generated in Isaac Lab on the Mug in Drawer task defined in mindmap.
The task was created to evaluate robot manipulation policies on their spatial memory capabilities.
With this (partial) dataset you can generate the full dataset used for mindmap model training,
run a mindmap training or evaluate mindmap open/closed loop.
This dataset is for research and development only.
Dataset Owner(s):… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-mindmap-Franka-Mug-in-Drawer.Physical-AI-AV-FR
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,909 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-FR.Physical-AI-AV-ES
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,674 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-ES.Physical-AI-AV-DE
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 324,105 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 10 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-DE.Physical-AI-AV-IT
PhysicalAI-AV-SFT
Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language
waypoint-prediction model. Contains 29,991 samples from 150 000
driving scenes (18 seconds per scene, sampled at anchor times 2s..16s) recorded in the
United States.
Format
WebDataset — 3 uncompressed .tar shards,
each containing pairs of files per sample:
Entry
Description
{key}.jpg
Front-facing wide-angle camera frame (JPEG quality 95, 640 × 360 px)
{key}.json… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-IT.
