datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hybrid_Neural_World_Models
Hybrid Neural World Models
Training, validation, test, and out-of-distribution (OOD) trajectories for the
three physical systems used in Hybrid Neural World Models (Pranav Lakshmanan,
Paras Chopra). The accompanying code, checkpoints, and paper define a single
neural surrogate that predicts states at any horizon plus a step-doubling
trust signal that flags when its forecasts can be trusted.
This repo contains the raw trajectory data only. Models / training code live
separately.… See the full description on the dataset page: https://huggingface.co/datasets/PraLak/Hybrid_Neural_World_Models.repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
sparse_world_modelsso101-task2-720p-whole-arm-v3-cleanedThis dataset was created using LeRobot.
Dataset Description
Recovered and cleaned SO-101 task 2 dataset. Bad final source episodes 95 and 96 were removed. V3 data, episode metadata, and video shard indices are compact and contiguous.
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 95,
"total_frames": 73988,
"total_tasks": 1,
"chunks_size": 1000… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v3-cleaned.world-models-eval
DreamGrasp: Processed LIBERO Manipulation Demonstrations
Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one?
This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides:
Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.browser-world-models-transitions
Browser World Models — Transitions
(before screenshot, action, after screenshot) transitions from real websites, for training and
evaluating a world model ("simulator") that predicts the consequence of a web action. Part of
the browser-world-models project.
How it was collected
An LLM policy (gpt-5.4-mini) drove vercel-labs/agent-browser
over the WebVoyager task set (642 tasks / 15 sites); the
task questions are the goals. For every action we saved a screenshot… See the full description on the dataset page: https://huggingface.co/datasets/sudac/browser-world-models-transitions.browser-world-models-webarena-reddit
browser-world-models: WebArena reddit transitions
Browser state transitions collected by running the 106 official WebArena reddit tasks
against a self-hosted, reproducible WebArena (Postmill) instance (rootless Apptainer on
SLURM). Each row: before/after screenshots, the Set-of-Marks action between them, the task
goal, and LLM-judge next-state labels in two target formats (free prose and a fixed
CHANGED/NAV/CONTENT/UI/ERROR template).
~726 transitions / 110 trajectories / 1… See the full description on the dataset page: https://huggingface.co/datasets/sudac/browser-world-models-webarena-reddit.demos
Demonstrations
We release a total of 4020 demonstrations across 200 tasks from MMBench.
RMFS_World_Models_Dataset
RMFS World Models — Speed Governor Corpora
Egocentric multi-camera trajectory data from a Robotic Mobile Fulfillment System (RMFS)
warehouse simulator, collected to train a World Model (VAE → MDN-RNN → CMA-ES controller,
after Ha & Schmidhuber 2018) that acts as a per-robot speed governor.
Research project at NTUST, Prof. Chou's CITI lab, extending the
RAWSim-O discrete-event simulator.
The controller does not choose where a robot goes. A* owns routing. The controller chooses… See the full description on the dataset page: https://huggingface.co/datasets/EthanGlueck/RMFS_World_Models_Dataset.so101-task2-720p-whole-arm-v4-freshThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 38664,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v4-fresh.spatial_world_modelsso101-task2-720p-whole-arm-v3-cleaned-trimmedThis dataset was created using LeRobot.
Dataset Description
Recovered and cleaned SO-101 task 2 dataset. Bad final source episodes 95 and 96 were removed. V3 data, episode metadata, and video shard indices are compact and contiguous.
License: apache-2.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 95,
"total_frames": 73988,
"total_tasks": 1,
"chunks_size": 1000… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v3-cleaned-trimmed.so101-task2-720p-whole-arm-v4-fresh-trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 38664,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task2-720p-whole-arm-v4-fresh-trimmed.so101-task1-720p-whole-arm-trimmed-subsampled-10fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0","robot_type": "so_follower",
"total_episodes": 201,
"total_frames": 62169,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task1-720p-whole-arm-trimmed-subsampled-10fps.inference-viz-1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 60,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/inference-viz-1.so101-task1-720p-whole-arm-trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 201,
"total_frames": 62169,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:201"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rl26-world-models/so101-task1-720p-whole-arm-trimmed.aat-real-world
Real-World AAT Materials Dataset
This dataset contains 189,523 real-world examples of cultural heritage object material descriptions paired with their corresponding Art & Architecture Thesaurus (AAT) material classifications.
Dataset Description
The dataset is designed for training models to extract material information from cultural heritage object descriptions. Each example consists of:
Input: A real material description from cultural heritage collections
Output:… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/aat-real-world.so101-task2-720p-whole-arm-trimmedso101-task2-720p-whole-arm-cube-trimmedso101-task2-720p-whole-arm-v3-cleaned-trimmed-subsampled-10fpsWorld_model_sim
Datasets
The recorded SO-100 teleop data is not stored in git. It lives on the
Hugging Face Hub:
https://huggingface.co/datasets/PrajnaYang/World_model_sim
Download
pip install "huggingface_hub[cli]"
# fetch everything into ./datasets/
hf download PrajnaYang/World_model_sim --repo-type dataset --local-dir datasets
Layout
datasets/
├── raw/ recorded teleop episodes (per task)
│ ├── single_cube/{teleop,dagger}
│ └── multi_cube/teleop
└──… See the full description on the dataset page: https://huggingface.co/datasets/PrajnaYang/World_model_sim.so101-task1-720p-whole-arm-cube-trimmedso101-720p-whole-arm-side-to-side-trimmedso101-task2-720p-whole-arm-trimmed-subsampled-10fpsworld-models-data-newso101-task2-720p-whole-arm-v4-fresh-trimmed-subsampled-10fps
