datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
watercolour-rollouts-judge-led
Watercolour rollouts, judge-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 861 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the run with the original reward mix from the write-up, where the pairwise judge and
its hand-rated pool carry most… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-judge-led.watercolour-rollouts-hps-only
Watercolour rollouts, HPS-only run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 470 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step.
The point of the dataset is that it holds the whole run, not the good bits. Step 0 and
step 59 are both here, with the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-hps-only.watercolour-rollouts-hps-led
Watercolour rollouts, hps-led run
Browse these paintings in the gallery Space, by step and by reward, with the sketch that made each one.
Every rollout from a GRPO run that taught Qwen/Qwen3.5-35B-A3B to paint watercolours by
writing p5.brush sketches. 872 paintings, the
sketch that produced each one, and the reward it earned, indexed by training step. This
is the middle point of the project's three reward mixes: the generic preference model
holds most of the weight, the… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/watercolour-rollouts-hps-led.MobileWorld-ProgRM-Rollouts
MobileWorld ProgRM Rollouts
Private training release containing multimodal ProgRM and ORM data derived from
MobileWorld rollouts. The release is intended for reward-model training and
subsequent reinforcement-learning research.
Release
Version: v1.0.0
ProgRM code commit: 875ba94c74a9742e903acaa81895f5685de7eac1 on branch mobileworld
Tasks: 161
Canonical trajectories: 966
Splits: {"test": 161, "train": 644, "val": 161}
Screenshot shards: 11
Run… See the full description on the dataset page: https://huggingface.co/datasets/sharryXR/MobileWorld-ProgRM-Rollouts.
