arianraje/qwen3-4b-gdn-hybrid-init-dtfix
qwen3-4b-gdn-hybrid-init-dtfix
Surgery/init checkpoint (dt_bias fixed) — uniform 1:4 GDN conversion of Qwen3-4B with plain weight inheritance, dtbias from the fla reference init (surgeryreport.json gdn_decay_init; gdndecaystats.json = activation-level retention on wikitext: 0/864 dead heads vs 281/864 for the original series).
Part of a study converting full-attention Qwen3-4B into a GDN (gated DeltaNet) hybrid (27 of 36 layers converted, uniform 1:4 retention) and recovering capability via staged distillation. Checkpoints are stock Qwen3NextForCausalLM — load with AutoModelForCausalLM (transformers >= 4.57).
`-dtfix` series. transformers' Qwen3Next port initializes the GDN time-step bias with dt_bias.fill_(1.0); with A_log = log U(0,16) and the decay g = -A * softplus(a + dt_bias) that gives retention exp(-dt*A) < 1e-3 for 2/3 of heads at init — the recurrent state starts dead and the gradient into dt_bias/A_log is too small to recover it (the original qwen3-4b-gdn-hybrid-* series carries this through every stage). This series re-runs the pipeline from a surgery that uses the fla GatedDeltaNet reference init (dt ~ LogUniform(1e-3, 0.1), dt_bias = softplus^-1(dt); median retention 0.94). Everything else (inheritance map, data recipe, trainer, LRs, budgets) is unchanged.
