laion/a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B
a3-rl-laion_nemotron-gym-agent-workplace-v2-5-8B
RL (GRPO / rloon) checkpoint, **globalstep5**, trained from `laion/GLM-47-swesmith-sandboxes-withtests-oracleverified_120s-maxeps-131k-fixthink (Qwen3-8B architecture) on open-athena/nemotron-gym-agent-workplace-v2`.
a3 series #21 — small, data-limited run. The dataset is only 297 tasks, so 2 epochs exhausted at ~step 9 (NOT the max_steps=80 budget). Checkpoint global_step_5 was selected by the highest 5-period EMA (alpha=1/3) of reward/avg_raw_reward (EMA@5 = 0.3987, raw reward@5 = 0.4746). The step-9 export emitted no logged reward and the trailing trajectory was flat/declining (s1..s8: 0.541, 0.479, 0.287, 0.197, 0.475, 0.355, 0.357, 0.252). This is a legitimate but weak a3 ablation data point, not a strong checkpoint.
Training Traces
Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: [open-athena/a3-rl-laion_nemotron-gym-agent-workplace-v2](https://huggingface.co/datasets/open-athena/a3-rl-laion_nemotron-gym-agent-workplace-v2)
The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) — the same rollouts the policy was trained on after rollback / truncation.
