semianalysisai/Qwen3.6-35B-A3B-graph-color-GRPO-20260909
028
Qwen3.6-35B-A3B · Graph Coloring
A specialist checkpoint for graph-coloring puzzles, trained from Qwen3.6-35B-A3B with 40 updates of verifier-reward GRPO (Group Relative Policy Optimization).
Training
Evaluation and scope
- Development evaluation: 512 puzzles per domain. Scores are not reported in this model card.
- Held-out test set: not evaluated.
- Export: trained language-model weights. Multimodal capabilities are not validated.
- Validate independently before downstream use. Equivalence to the results reported in the Miles PR has not been established.
<details> <summary><strong>Reproducibility and implementation details</strong></summary>
Each specialist was trained on its own puzzle domain. Reasoning Gym generated 10,000 training puzzles per domain.
SGLang used the Triton MoE backend because automatic backend selection failed during weight transfer.
</details>
Companion checkpoint: Countdown
