datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
memory-rolloutswatercolour-rollouts-judge-led
Watercolour rollouts, judge-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 861 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the run with the original reward mix from the write-up, where the pairwise judge and
its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.watercolour-rollouts-hps-only
Watercolour rollouts, HPS-only run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 470 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step.
The point of the dataset is that it holds the whole run, not the good bits. Step 0 and
step 59 are both here, with the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-hps-only.screen-highlighter-4b-highlight-only-v1-rolloutsdipper_rolloutsastra-robodojo-rollouts
Astra RoboDojo Evaluation Records
Rollout records from the evaluations in GPT 6 Astra as an Embodied Policy,
by Jiayi Su, Yixin Zheng, Mi Yan, Li Yi, Zhizheng Zhang, and He Wang. This archive
includes action proposals, executed actions, observations, robot states,
model-provided explanations and reasoning summaries, and metadata for reproducing
the evaluation settings, together with a reader and documentation.
Report
Public controller source
Data schema and alignment… See the full description on the dataset page: https://huggingface.co/datasets/YuMoool/astra-robodojo-rollouts.watercolour-rollouts-hps-led
Watercolour rollouts, hps-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 872 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the middle point of the project's three reward mixes: the generic preference model
holds most of the weight, the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-hps-led.ur5e_pi05_rollouts
UR5e pi0.5 rollouts
Policy rollouts of mahgoobi/ur5e_pi05_10k on a
real UR5e with a Robotiq gripper, task "put the cup in the bowl", collected 2026-09-01 from
the deployment page in RoboResearch
(roboresearch.evaluation.ur5e.ui). 25 episodes at 20 Hz, action chunks of 50 steps
with action_steps of them executed per chunk (25 for every episode but the first, which
ran 10). 20 successes, 5 failures. Every failure is in a scene with
three or more distractors; the clean and… See the full description on the dataset page: https://huggingface.co/datasets/mahgoobi/ur5e_pi05_rollouts.screen-highlighter-2b-highlight-only-v1-rolloutsscreen-highlighter-2b-v1-rolloutsprofile-rollouts-v2lehome-dataset-rolloutsrollouts-vla0-500ep-naive-ghost-traceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 100182,
"total_tasks": 12,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/end-effector-trace-conditioning/rollouts-vla0-500ep-naive-ghost-trace.rollouts-vla0-500ep-naive-5frameThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 60,
"total_frames": 102010,
"total_tasks": 12,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/end-effector-trace-conditioning/rollouts-vla0-500ep-naive-5frame.rollouts_po2_train_2_cleaned_34521libero-branch-rollouts
LIBERO branch-rollout benchmark (EXP-20260722-004, C-line + Tier-1)
Data generated for a study on metric readiness for video world models: specifically,
whether a learned model's outputs actually respond to the action it is conditioned on,
rather than to the presence of a conditioning signal at all.
Everything here was produced by live MuJoCo simulation on an AMD MI325X node.
It is not a repackaging of the public LIBERO release. See "What is NOT here".
What is here… See the full description on the dataset page: https://huggingface.co/datasets/Minhao-VWM-metrics/libero-branch-rollouts.rollouts_po2_trainui_design_reasoner_rollouts_nemotronvlrollouts_po2_train_2ui-reasoning-effort-rollouts-full
UI Reasoning Effort Rollouts
Completed rollout records with reasoning traces, UIClip scores, token counts, latency, and source metadata.
Hub repo: https://huggingface.co/datasets/ciderlab/ui-reasoning-effort-rollouts-full
Rows: 667
train: 667
The dataset is stored with Hugging Face's native save_to_disk() format and embeds raw responses, HTML, screenshots, and prompt templates.
Columns
model_name: model or deployment name used for the rollout.
served_model:… See the full description on the dataset page: https://huggingface.co/datasets/ciderlab/ui-reasoning-effort-rollouts-full.rollouts_po2_val
