datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.PhysicalAI-VANTAGE-Bench
VANTAGE-BENCH
Video ANalysis Tasks Across Generalized Environments
Paper: VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models
Dataset Description
VANTAGE-BENCH is the first public benchmark purpose-built for evaluating visual understanding on video captured by fixed infrastructure cameras. It spans three real-world domains — warehouse, smart city / Intelligent Transportation Systems (ITS), and smart spaces — across six spatio-temporal… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-VANTAGE-Bench.PhysicalAI-ADE-US
PhysicalAI-US-ADE
Dataset Summary
PhysicalAI-US-ADE contains per-sample evaluation outputs for autonomous driving waypoint prediction on the US subset of the PhysicalAI NVIDIA dataset.
This dataset stores inference-time predictions and evaluation statistics for models evaluated on the dataset, organized by model name at the top level. Each model directory contains sample-level records for that model’s predictions against ground truth.
The current release includes… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-ADE-US.PhysicalAI-ADE-DEPhysicalAI-US-Evaluation
PhysicalAI-US-Evaluation
A held-out US evaluation set for the navigation planner: 19,744 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn — uniformly at random — from the pool of US scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-US-Evaluation.Physical_AIPhysicalAI-DE-Evaluation
PhysicalAI-DE-Evaluation
A held-out German evaluation set for the navigation planner: 19,999 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn from the pool of German scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT) stage, and
the… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-DE-Evaluation.tanitad-physicalai-w120-256x640cyl
TanitAD PhysicalAI w120 256x640 cylindrical caches
Derived from NVIDIA PhysicalAI-Autonomous-Vehicles. Not an original dataset.
physicalai-train-e438721ae894-w120-256x640cyl - canonical train corpus, parity key
e438721ae894, skip-hash f09e44db. Parity is sacred: anything that re-selects episodes
breaks cross-arm comparability.
physicalai-val-0c5f7dac3b11-w120-256x640cyl - matching val corpus.
Geometry: --frame-h 256 --frame-w 640 --frame-hfov 120 --projection cylindrical… See the full description on the dataset page: https://huggingface.co/datasets/Sayood/tanitad-physicalai-w120-256x640cyl.
