datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
Dataset Description:
PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.code-world-model-project-page-videos
Code World Model Project Page Videos
Public research-demo video assets used by the Code World Model project page.
The gallery/ directory contains aligned RGB and proxy videos for interactive comparison.
PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
PhysicalAI SDG-Warehouse
PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
Dataset Description:
The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research.
Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card
Dataset Description
PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.phyground
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Project page ·
Paper ·
Evaluation code ·
PhyJudge-9B
PhyGround is a criteria-grounded benchmark for diagnosing physical failures in
generated video. It contains 250 prompts covering 13 observable physical
laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is
paired with a first-frame image, 10 released generation configurations, and
applicable-law labels.
The Hub repository includes:… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.world_modelaction-world-model-atlas-1500-media-20260914
Action World Model Atlas
Public browsing previews for 1,500 unique action clips from the completed
6,033-video bundle. OpenPixel2Play, Gaming 500 Hours, and Xiaoluo each contribute
500 examples. All 46 games in the completed bundle are represented.
Videos preserve the full five-second duration and 81 frames. They are existing
browser previews and can be smaller than the native training videos. Video and
poster checksums are verified against the source media manifests.… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action-world-model-atlas-1500-media-20260914.world-model-physics
Rapidata Physics Benchmark
Built by Rapidata.
Do video and world models understand physics? We gave 25 video- and world models the same
real-world starting frame and scene description from Physics-IQ and asked
them to predict what happens next. ~283,000 human votes, collected with the
Rapidata Python SDK, decided which continuation is more realistic — with the
real recording competing as a hidden 26th participant.
Each row is a head-to-head matchup between two participants on… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/world-model-physics.asl-trunk-world-model
ASL soft trunk robot: world-model dataset and training package
Research data of the Stanford ASL trunk robot project (recorded 2026-09-10/11). Code: https://github.com/mellowyellow71/trunk-stack (branch world-model, a fork of StanfordASL/trunk-stack). Everything a new person needs to reproduce or continue the work is in this repository plus that branch; no access to the original machines is required.
What is here
path
content… See the full description on the dataset page: https://huggingface.co/datasets/melloyello/asl-trunk-world-model.WorldModelBenchgame-world-model-dataset-visual-guideaction-world-model-sample-300
Action World Model: Random 300-case samples
This public index contains deterministic random samples of 300 cases from each source dataset (seed: 20260918). The JSONL files preserve source identifiers, source paths/URLs, and local paths where available. Video files are not duplicated in this index.
Datasets: ABot-World-Explorer-500h, game-recordings-v3, EmbodiedWorld-200K.
world_model_data_ours_v3world_model_real_rollout_gennms_hitl_world_model
No Man's Sky High-Fidelity Human-in-the-loop World Model Dataset
Overview
This dataset is designed for world model training using real human gameplay data from No Man’s Sky.It captures high-fidelity human–computer interaction by recording both the game video and time-aligned input actions, preserving the realistic latency characteristics of a human-in-the-loop system.
Compared with “internal game state” datasets, this dataset retains the physical interaction chain (input… See the full description on the dataset page: https://huggingface.co/datasets/HuberyLL/nms_hitl_world_model.lerobot_pick_and_place_dataset_world_modelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 30,
"total_frames": 13572,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Clementppr/lerobot_pick_and_place_dataset_world_model.world-model-gameplay-recording
World Model Gameplay Recording
Action-conditioned gameplay video dataset for world model training. Contains synchronized high-resolution gameplay recordings with frame-accurate input action logs (gamepad, keyboard, mouse) from multiple AAA game titles.
Games Included
#
Game
Session ID
Duration
Video Format
Video Size
Input Type
1
Game Session 1
fwa0NekU
~5 min
MKV
582 MB
Gamepad + Keyboard
2
The Legend of Zelda: Tears of the Kingdom
g4qz1DLq
~15 min
MP4… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/world-model-gameplay-recording.so101-task2-720p-whole-arm-v3-cleanedThis dataset was created using LeRobot.
Dataset Description
Recovered and cleaned SO-101 task 2 dataset. Bad final source episodes 95 and 96 were removed. V3 data, episode metadata, and video shard indices are compact and contiguous.
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 95,
"total_frames": 73988,
"total_tasks": 1,
"chunks_size": 1000… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v3-cleaned.world-models-eval
DreamGrasp: Processed LIBERO Manipulation Demonstrations
Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one?
This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides:
Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.world_modelso101-task2-720p-whole-arm-v4-freshThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 38664,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v4-fresh.WorldModelAffordanceso101-task2-720p-whole-arm-v4-fresh-trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 38664,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v4-fresh-trimmed.so101-task2-720p-whole-arm-v3-cleaned-trimmedThis dataset was created using LeRobot.
Dataset Description
Recovered and cleaned SO-101 task 2 dataset. Bad final source episodes 95 and 96 were removed. V3 data, episode metadata, and video shard indices are compact and contiguous.
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 95,
"total_frames": 73988,
"total_tasks": 1,
"chunks_size": 1000… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v3-cleaned-trimmed.so101-task2-720p-whole-arm-cube-trimmedso101-task1-720p-whole-arm-trimmed-subsampled-10fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 201,
"total_frames": 62169,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task1-720p-whole-arm-trimmed-subsampled-10fps.inference-viz-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 60,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/inference-viz-1.SO101-world-model-5fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 107,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 5,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/julie-trrsr/SO101-world-model-5fps.so101-task1-720p-whole-arm-trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 201,
"total_frames": 62169,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task1-720p-whole-arm-trimmed.
