CoolFace
Modelpublic

joey00072/Qwen3.6-35B-A3B-countdown-GRPO-20260909

sourceHugging Faceupdated 18d agoView on Hugging Face
0likes18downloads
Model Card

Puzzle specialist teacher

This checkpoint specializes in countdown puzzles through 40 verifier-reward GRPO updates. The base model is Qwen/Qwen3.6-35B-A3B at revision 995ad96eacd98c81ed38be0c5b274b04031597b0. Miles source is radixark/miles PR 3116 at revision 8f8e4dff55d80bd9d36c3490c333a119f38480ed. Reasoning Gym revision 49b07130b3fcd12f2d064bba7c43869543a0e7e7 generated 10,000 training puzzles per domain. Each update used 32 prompts with 8 responses each, a global batch of 256, learning rate 1e-6, and a 256-token response cap with thinking disabled. Each teacher trained on one node with eight NVIDIA B200 GPUs. SGLang used the Triton MoE backend because automatic backend selection failed during weight transfer. The export contains the trained language-model weights. No claim is made about multimodal capabilities or equivalence to the PR's reported results. Development evaluations used 512 puzzles per domain. The held-out test set has not been evaluated. The model requires independent validation before downstream use.