CoolFace
Datasetpublic

penfever/jupiter-tasktrove-dapo-artifacts

Jupiter TaskTrove DAPO — campaign artifacts Standalone artifact archive for the jupiter-tasktrove-dapo campaign: a six-arm objective ablation (GRPO control vs DAPO variants vs GSPO) training Qwen/Qwen3-Coder-30B-A3B-Instruct with agentic RL (MarinSkyRL / SkyRL fully-async GRPO trainers, Harbor + Daytona sandboxed terminus-2 rollouts) on competitive-programming tasks, run on JSC Jupiter (GH200) 2026-08-20 → 2026-08-31. The campaign closed inconclusive (platform degradation, a… See the full description on the dataset page: https://huggingface.co/datasets/penfever/jupiter-tasktrove-dapo-artifacts.

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
0likes156downloads
Dataset Card

Jupiter TaskTrove DAPO — campaign artifacts

Standalone artifact archive for the jupiter-tasktrove-dapo campaign: a six-arm objective ablation (GRPO control vs DAPO variants vs GSPO) training Qwen/Qwen3-Coder-30B-A3B-Instruct with agentic RL (MarinSkyRL / SkyRL fully-async GRPO trainers, Harbor + Daytona sandboxed terminus-2 rollouts) on competitive-programming tasks, run on JSC Jupiter (GH200) 2026-08-20 → 2026-08-31.

The campaign closed inconclusive (platform degradation, a mid-experiment dataset quality issue, and a DAPO starvation failure mode) — see the linked experiment issue for the human summary and the operating policy.

Contents (artifacts/)

pathwhat it is
configs/the six arm launch YAMLs (d0 GRPO, d1 DAPO-core, d2 full DAPO, d3 §3.4-only, d4 DAPO+partial-credit, d5 GSPO)
scripts/campaign launch/compose-check/gate-watch tooling used on Jupiter
curves/per-arm training-curve extraction (CSV cache + plots + plot_dapo_curves.py)
evidence/preserved incident evidence (e.g. the wandb.init shared-mode hang)
preserved/cleanup-run preservations: metrics reports/tables, export/upload logs for all six arms
retired/logs of retired/superseded chains
cleanup_jtd-*.mdfull rl-agentic-job-cleanup run logs (d0, d5, d1–d4)
EVAL_RESULTS.mdsealed final eval table (arms vs baseline)
GRPO_VS_DAPO.mdmechanical GRPO↔DAPO difference map with per-knob provenance
ASYNC_DAPO_IMPLEMENTATIONS.mdsurvey of public async-DAPO implementations
TASK_VARIANCE.md + task_variance_raw.jsonper-task pass-rate variance analysis

Published outputs

  • —Weights: laion/jtd-d0-69-30B, laion/jtd-d1-18-30B, laion/jtd-d2-20-30B, laion/jtd-d3-60-30B, laion/jtd-d4-34-30B, laion/jtd-d5-69-30B
  • —Training traces: penfever/jtd-d0 … penfever/jtd-d5 (d0/d5 full set; d1–d4 at a 25% conservation subsample)

Links

  • —Experiment issue (human tl;dr + operating policy): https://github.com/marin-community/marin/issues/8829
  • —Base model: Qwen/Qwen3-Coder-30B-A3B-Instruct