CoolFace
Modelpublic

arianraje/qwen3-4b-gdn-hybrid-stage3-200M-OPD-dtfix

sourceHugging Facemitupdated 23d agoView on Hugging Face
0likes18downloads
Model Card

qwen3-4b-gdn-hybrid-stage3-200M-OPD-dtfix

Model weights of the Qwen3-4B GDN-hybrid dt_bias-fixed line, stage-3 on-policy distillation (OPD), 200M generated tokens.

Run stage3-opd-dtfix-200M-v1 (W&B id qwen-wsd-flat-dtfix-200M-v1-20260903), final snapshot at step 1247 / 200,104,080 consumed generated tokens, horizon 16K, flat LR 2e-5 with a 150-step warmup and a 300-step (39.3M-token) linear decay to 1e-6.

Lineage

Student: arianraje/qwen3-4b-gdn-hybrid-stage2b-kd-dtfix (stage-2b KD on the dtbias-fixed stage-1 alignment `arianraje/qwen3-4b-gdn-hybrid-stage1-align-dtfix`). Teacher: `Qwen/Qwen3-4B`. This is a like-for-like rerun of the ORIGINAL `stage3-200M-OPD` rung of `pinkskin/qwen3-4b-gdn-wsd-ladder` (`wsd-flat-200M`, WSDexps/configs/wsdflat.yaml): same prompt mixture (stage3promptsv1), validation set, schedule, genbatch and queue factor. The only intended difference in the whole line is the GDN dt_bias initialisation: the original line inherited the transformers dt_bias.fill_(1.0) init (saturated decay gates); this line uses the fla-style init (dt_bias median -4.56).

Evals (this snapshot)

  • —Commonsense/MMLU (lm-eval, likelihood): piqa 73.5, hellaswag 62.0, arceasy 71.4, arcchallenge 48.1, winogrande 63.2, mmlu 58.7
  • —Math (n=8 samples, pass@1): nothinkgsm8k 80.8, nothinkmath500 66.7, think_math500 83.1, aime24 23.8, aime25 17.9

The result JSONs are included (likelihood_battery.json, reasoning_bench_*.json); the head-to-head against the original line is in artifacts/dual_family_eval_24h/TABLES_DTFIX.md of the project repo.

Custom GDN-hybrid architecture - register before loading (see project repo).

Source Git commit: d86fbef09d35f4e4d7943ec51d2b3732eb1fed46