datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card
Dataset Description
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding.
Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.world_modelshowui-worldmodel-results
ShowUI + WorldModel: Training Results & Artifacts
This dataset contains evaluation results, world models, and training data from the ShowUI + WorldModel integration project.
📦 Contents
1. Evaluation Results (results/miniwob_predictions/)
Size: ~460MB
Format: JSONL files with episode-level predictions
Tasks: 9 MiniWoB++ tasks evaluated with ShowUI agent
Includes:
Task success/failure outcomes
Action predictions and execution traces
World model… See the full description on the dataset page: https://huggingface.co/datasets/zhongweixie/showui-worldmodel-results.action-world-model-atlas-1500-media-20260914
Action World Model Atlas
Public browsing previews for 1,500 unique action clips from the completed
6,033-video bundle. OpenPixel2Play, Gaming 500 Hours, and Xiaoluo each contribute
500 examples. All 46 games in the completed bundle are represented.
Videos preserve the full five-second duration and 81 frames. They are existing
browser previews and can be smaller than the native training videos. Video and
poster checksums are verified against the source media manifests.… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action-world-model-atlas-1500-media-20260914.Mobile-GUI-Worldmodel-SFT
Mobile-GUI-Worldmodel-SFT
This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text.
Repository Layout
.
├── GUI-agent-main/ # Data annotation scripts and examples
├── eval/ # Evaluation assets
│ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.menti-bench
Menti-Bench
Menti-Bench is a manually constructed, quality-controlled benchmark of situated decision scenarios for evaluating Mental World Modeling (MWM): whether a model can predict what a target agent will actually do next, in scenes where the correct prediction depends on tracking each agent's beliefs, knowledge access, goals, emotions, and social constraints rather than the physical scene alone.
Each instance presents a short story (text, an image sequence, or a sounding… See the full description on the dataset page: https://huggingface.co/datasets/mental-world-model/menti-bench.WorldModelBenchgame-world-model-dataset-visual-guideSyn4D_worldmodel_benchmarkvjepa2-reacher-world-modelbrowser-world-models-transitions
Browser World Models — Transitions
(before screenshot, action, after screenshot) transitions from real websites, for training and
evaluating a world model ("simulator") that predicts the consequence of a web action. Part of
the browser-world-models project.
How it was collected
An LLM policy (gpt-5.4-mini) drove vercel-labs/agent-browser
over the WebVoyager task set (642 tasks / 15 sites); the
task questions are the goals. For every action we saved a screenshot… See the full description on the dataset page: https://huggingface.co/datasets/sudac/browser-world-models-transitions.WorldModelAffordancebrowser-world-models-webarena-reddit
browser-world-models: WebArena reddit transitions
Browser state transitions collected by running the 106 official WebArena reddit tasks
against a self-hosted, reproducible WebArena (Postmill) instance (rootless Apptainer on
SLURM). Each row: before/after screenshots, the Set-of-Marks action between them, the task
goal, and LLM-judge next-state labels in two target formats (free prose and a fixed
CHANGED/NAV/CONTENT/UI/ERROR template).
~726 transitions / 110 trajectories / 1… See the full description on the dataset page: https://huggingface.co/datasets/sudac/browser-world-models-webarena-reddit.worldmodel_wa_multimodal_dummy
Dataset Card for "worldmodel_wa_multimodal_dummy"
More Information needed
