laion/a3-rl-laion_nemotron-gym-instruction-following-structured-75-8B
a3-rl-laionnemotron-gym-instruction-following-structured (RL, globalstep 75)
RL fine-tune (SkyRL) of laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink (Qwen3-8B) on open-athena/nemotron-gym-instruction-following-structured.
Checkpoint selected by 5-period EMA (alpha=1/3) of reward/avg_raw_reward over the full 80-step trajectory (reconstructed across 11 chain-restart .out logs). global_step 75 is the highest-EMA saved export (EMA=0.9512, step reward=0.9902). The run trained the full max_steps=80 (final reward ~0.92, pass@8 ~0.95), surviving ~5 vLLM crash-restarts.
Training Traces
Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: [open-athena/a3-rl-laion_nemotron-gym-instruction-following-structured](https://huggingface.co/datasets/open-athena/a3-rl-laion_nemotron-gym-instruction-following-structured)
The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) — the same rollouts the policy was trained on after rollback / truncation.
