CoolFace
Modelpublic

semianalysisai/Qwen3.6-35B-A3B-graph-color-GRPO-20260909

sourceHugging Faceupdated 18d agoView on Hugging Face
0likes28downloads
Model Card

Qwen3.6-35B-A3B · Graph Coloring

A specialist checkpoint for graph-coloring puzzles, trained from Qwen3.6-35B-A3B with 40 updates of verifier-reward GRPO (Group Relative Policy Optimization).

Training

SettingValue
Base modelQwen/Qwen3.6-35B-A3B
Training puzzles10,000, generated with Reasoning Gym
GRPO updates40
Batch per update32 prompts × 8 responses = 256 responses
Learning rate1e-6
Response limit256 tokens; thinking disabled
Hardware1 node with 8 × NVIDIA B200 GPUs
Training frameworkMiles

Evaluation and scope

  • —Development evaluation: 512 puzzles per domain. Scores are not reported in this model card.
  • —Held-out test set: not evaluated.
  • —Export: trained language-model weights. Multimodal capabilities are not validated.
  • —Validate independently before downstream use. Equivalence to the results reported in the Miles PR has not been established.

<details> <summary><strong>Reproducibility and implementation details</strong></summary>

Each specialist was trained on its own puzzle domain. Reasoning Gym generated 10,000 training puzzles per domain.

ComponentSource revision
Base model995ad96eacd98c81ed38be0c5b274b04031597b0
Miles — radixark/miles, PR #31168f8e4dff55d80bd9d36c3490c333a119f38480ed
Reasoning Gym49b07130b3fcd12f2d064bba7c43869543a0e7e7

SGLang used the Triton MoE backend because automatic backend selection failed during weight transfer.

</details>


Companion checkpoint: Countdown