penfever/jupiter-tasktrove-dapo-artifacts
Jupiter TaskTrove DAPO — campaign artifacts Standalone artifact archive for the jupiter-tasktrove-dapo campaign: a six-arm objective ablation (GRPO control vs DAPO variants vs GSPO) training Qwen/Qwen3-Coder-30B-A3B-Instruct with agentic RL (MarinSkyRL / SkyRL fully-async GRPO trainers, Harbor + Daytona sandboxed terminus-2 rollouts) on competitive-programming tasks, run on JSC Jupiter (GH200) 2026-08-20 → 2026-08-31. The campaign closed inconclusive (platform degradation, a… See the full description on the dataset page: https://huggingface.co/datasets/penfever/jupiter-tasktrove-dapo-artifacts.
Jupiter TaskTrove DAPO — campaign artifacts
Standalone artifact archive for the jupiter-tasktrove-dapo campaign: a six-arm objective ablation (GRPO control vs DAPO variants vs GSPO) training Qwen/Qwen3-Coder-30B-A3B-Instruct with agentic RL (MarinSkyRL / SkyRL fully-async GRPO trainers, Harbor + Daytona sandboxed terminus-2 rollouts) on competitive-programming tasks, run on JSC Jupiter (GH200) 2026-08-20 → 2026-08-31.
The campaign closed inconclusive (platform degradation, a mid-experiment dataset quality issue, and a DAPO starvation failure mode) — see the linked experiment issue for the human summary and the operating policy.
Contents (artifacts/)
Published outputs
- Weights:
laion/jtd-d0-69-30B,laion/jtd-d1-18-30B,laion/jtd-d2-20-30B,laion/jtd-d3-60-30B,laion/jtd-d4-34-30B,laion/jtd-d5-69-30B - Training traces:
penfever/jtd-d0…penfever/jtd-d5(d0/d5 full set; d1–d4 at a 25% conservation subsample)
Links
- Experiment issue (human tl;dr + operating policy): https://github.com/marin-community/marin/issues/8829
- Base model: Qwen/Qwen3-Coder-30B-A3B-Instruct
