laion/a3-rl-laion_nemotron-gym-identity-following-v2-65-8B
043
a3-rl-laion_nemotron-gym-identity-following-v2 (step 65, 8B)
RL (GRPO/rloo_n) finetune trained with SkyRL on Jupiter (a3 series #22).
- Base model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
- Training dataset: open-athena/nemotron-gym-identity-following-v2
- Checkpoint selection: globalstep 65, chosen by highest 5-period EMA (alpha=1/3) of `reward/avgrawreward` across the full 80-step run (EMA=0.9996, raw reward at step=1.0000). Run completed its full maxsteps=80 at near-perfect reward.
Training Traces
Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: [open-athena/a3-rl-laion_nemotron-gym-identity-following-v2](https://huggingface.co/datasets/open-athena/a3-rl-laion_nemotron-gym-identity-following-v2)
The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) the same rollouts the policy was trained on after rollback / truncation.
