datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.code-world-model-inference-examples-40
Inference examples
This directory contains 40 numbered, independent inference examples.
Every example uses only its public number; source case names and internal paths
are intentionally omitted.
Each numbered directory contains:
first_frame.png: exact 1536x864 generated RGB first frame used by inference.
prompts/*.txt: the exact rolling long-inference prompts used for the result.
condition/*.npz: ordered lossless condition chunks.
metadata.json: frame count, FPS, prompt windows… See the full description on the dataset page: https://huggingface.co/datasets/NTU-yiwen/code-world-model-inference-examples-40.financial-world-model
Twelve Data World Model Dataset
A multi-modal financial time-series dataset built from Twelve Data
market data. Each timeframe is published in three parallel views:
bars_* — OHLCV bars enriched with causal technical indicators and macro
context, in Parquet.
text_* — instruction-tuning prompts/labels derived from the bars, in
JSONL.
trajectories_* — fixed-length rolling windows of state vectors plus
next-state pairs, suitable for world-model / sequence-model training, in… See the full description on the dataset page: https://huggingface.co/datasets/twelvedata/financial-world-model.phyground
PhyGround: Benchmarking Physical Reasoning in Generative World Models
Project page ·
Paper ·
Evaluation code ·
PhyJudge-9B
PhyGround is a criteria-grounded benchmark for diagnosing physical failures in
generated video. It contains 250 prompts covering 13 observable physical
laws across solid-body mechanics, fluid dynamics, and optics. Each prompt is
paired with a first-frame image, 10 released generation configurations, and
applicable-law labels.
The Hub repository includes:… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.world-model-physics
Rapidata Physics Benchmark
Built by Rapidata.
Do video and world models understand physics? We gave 25 video- and world models the same
real-world starting frame and scene description from Physics-IQ and asked
them to predict what happens next. ~283,000 human votes, collected with the
Rapidata Python SDK, decided which continuation is more realistic — with the
real recording competing as a hidden 26th participant.
Each row is a head-to-head matchup between two participants on… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/world-model-physics.world_model_corpus
Dataset Card for World Model Corpus
Paper | GitHub
The world model corpus contains a set of generated trajectories that are shaped for text-based world modeling task as used by the paper: "Masked Diffusion Language Models are Strong and
Steerable Text-Based World Models for Agentic RL". The dataset contains trajectories from nine distinct environments: Tau2Bench, SWE-Smith, DeepresearchQA, Openresearcher, Gorilla/BFCLv4, Webshop, Toolathlon, Pandora and Coderforge.… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/world_model_corpus.World-Model
🌍 World Model Bench (WM Bench) v1.0
Beyond FID — Measuring Intelligence, Not Just Motion
WM Bench is the world's first benchmark for evaluating the cognitive capabilities of World Models and Embodied AI systems.
🎯 Why WM Bench?
Existing world model evaluations focus on:
FID / FVD — image and video quality ("Does it look real?")
Atari scores — performance in fixed game environments
WM Bench measures something different: Does the model think correctly?… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/World-Model.Mobile-GUI-Worldmodel-SFT
Mobile-GUI-Worldmodel-SFT
This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text.
Repository Layout
.
├── GUI-agent-main/ # Data annotation scripts and examples
├── eval/ # Evaluation assets
│ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.bytesized32-world-model-cotSee https://github.com/thuml/RLVR-World for examples for using this dataset.
Citation
@article{wu2025rlvr,
title={RLVR-World: Training World Models with Reinforcement Learning},
author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long},
journal={arXiv preprint arXiv:2505.13934},
year={2025},
}
menti-bench
Menti-Bench
Menti-Bench is a manually constructed, quality-controlled benchmark of situated decision scenarios for evaluating Mental World Modeling (MWM): whether a model can predict what a target agent will actually do next, in scenes where the correct prediction depends on tracking each agent's beliefs, knowledge access, goals, emotions, and social constraints rather than the physical scene alone.
Each instance presents a short story (text, an image sequence, or a sounding… See the full description on the dataset page: https://huggingface.co/datasets/mental-world-model/menti-bench.game-world-model-dataset-visual-guiderepro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
reasoning_world_modelwebarena-world-model-cotSee https://github.com/thuml/RLVR-World for examples for using this dataset.
Citation
@article{wu2025rlvr,
title={RLVR-World: Training World Models with Reinforcement Learning},
author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long},
journal={arXiv preprint arXiv:2505.13934},
year={2025},
}
world_model_real_rollout_genworld_model_for_wa_desc_with_tao_dataset
Dataset Card for "world_model_for_wa_desc_with_tao_dataset"
More Information needed
LiteCoder-Terminal-World-Model-SFTworld-model-gameplay-recording
World Model Gameplay Recording
Action-conditioned gameplay video dataset for world model training. Contains synchronized high-resolution gameplay recordings with frame-accurate input action logs (gamepad, keyboard, mouse) from multiple AAA game titles.
Games Included
#
Game
Session ID
Duration
Video Format
Video Size
Input Type
1
Game Session 1
fwa0NekU
~5 min
MKV
582 MB
Gamepad + Keyboard
2
The Legend of Zelda: Tears of the Kingdom
g4qz1DLq
~15 min
MP4… See the full description on the dataset page: https://huggingface.co/datasets/obaydata/world-model-gameplay-recording.social-world-model-v6-qwen3.5-397B-clean-semdedupsubjective-world-model
Subjective World Model — Preview Dataset
by Atypica
Preview release. All records are illustrative archetypes constructed from real interview methodology — not records of specific individuals.
Introduction
This dataset trains AI agents to simulate how specific people think and make decisions — not statistical averages or fictional archetypes, but individuals with their own logic, habits, and contradictions.
Two record types:
persona — one individual: an AI-conducted… See the full description on the dataset page: https://huggingface.co/datasets/atypica/subjective-world-model.world_model_data
world_model_data
Recorded terminal-agent trajectories paired with two prediction targets: the original
abstracted observation, and a newer transition abstraction that also states the
command's outcome, its effects on the machine, and the agent's workflow phase.
Research artifact, not a released dataset. Everything below is read off the files in this
repo and the run logs. Anything that was not recorded is marked as such rather than guessed at.
See TARGETS.md for the full… See the full description on the dataset page: https://huggingface.co/datasets/joykirat/world_model_data.browser-world-models-transitions
Browser World Models — Transitions
(before screenshot, action, after screenshot) transitions from real websites, for training and
evaluating a world model ("simulator") that predicts the consequence of a web action. Part of
the browser-world-models project.
How it was collected
An LLM policy (gpt-5.4-mini) drove vercel-labs/agent-browser
over the WebVoyager task set (642 tasks / 15 sites); the
task questions are the goals. For every action we saved a screenshot… See the full description on the dataset page: https://huggingface.co/datasets/sudac/browser-world-models-transitions.world_model_for_wa_desc_with_tao_formatted_w_cot
Dataset Card for "world_model_for_wa_desc_with_tao_formatted_w_cot"
More Information needed
world-model-pi-sft-formatted
world-model-pi-sft-formatted
Role-flipped formatted SFT data for training a world model (environment simulator) over the
TeichAI pi coding-agent harness. Source traces: TeichAI/DeepSeek-v4-Pro-Agent (4006 sessions).
This dataset is stage 1 of a two-stage pipeline: stage 2 (world-model-pi-sft-with-cot)
injects synthetic <think> rationales into every assistant target via a reasoning-model
teacher; this stage 1 dataset has only the placeholder <COT_PLACEHOLDER> token that stage 2… See the full description on the dataset page: https://huggingface.co/datasets/kfallah/world-model-pi-sft-formatted.counterfactual-intervention-integrity-worldmodel-v01
Dataset
ClarusC64/counterfactual-intervention-integrity-worldmodel-v01
This dataset tests one capability.
Can a model reason cleanly about interventions without breaking the world.
Core rule
Interventions have local consequences.
Changing one thing
must change what depends on it
must not change what does not
No magic.
No silent propagation.
No ignored causes.
Canonical labels
WITHIN_SCOPE
OUT_OF_SCOPE
Files… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/counterfactual-intervention-integrity-worldmodel-v01.world_model_eval_logsworld_model_for_wa_qa_formatted
Dataset Card for "world_model_for_wa_qa_formatted"
More Information needed
browser-world-models-webarena-reddit
browser-world-models: WebArena reddit transitions
Browser state transitions collected by running the 106 official WebArena reddit tasks
against a self-hosted, reproducible WebArena (Postmill) instance (rootless Apptainer on
SLURM). Each row: before/after screenshots, the Set-of-Marks action between them, the task
goal, and LLM-judge next-state labels in two target formats (free prose and a fixed
CHANGED/NAV/CONTENT/UI/ERROR template).
~726 transitions / 110 trajectories / 1… See the full description on the dataset page: https://huggingface.co/datasets/sudac/browser-world-models-webarena-reddit.world_model_for_wa_description_formatted_w_cotstate-continuity-temporal-coherence-worldmodel-v01
Dataset
ClarusC64/state-continuity-temporal-coherence-worldmodel-v01
This dataset tests one capability.
Can a model preserve a coherent world state across time.
Core rule
The world has memory.
Once something changeslater descriptions must reflect that change.
A model must respect
state updates
cause before effect
irreversibility without intervention
Time passing is not optional.
Canonical labels
WITHIN_SCOPE
OUT_OF_SCOPE
Files… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/state-continuity-temporal-coherence-worldmodel-v01.
