nlproj/gated_deltanet_350M_rfull_bs256_lr3e-4_steps10000
082
GatedDeltaNet 350M (full rank) — Low-rank Fast-Weight Ablation
Pretrained 350M-parameter GatedDeltaNet with low-rank parameterization (rfull) on FineWeb-Edu. Part of a multi-cell ablation across 4 archs × {r32, r64, r256, rfull} (plus GDN extras r512) studying whether constraining the q/k/v fast-weight projections (or LaCT's SwiGLU MLP) to low rank can match or exceed full-rank performance at the 350M scale.
Training
Eval results
- FineWeb-Edu val PPL:
12.21 - LAMBADA acc: 0.315
- HellaSwag acc_norm: 0.401
- ARC-Easy acc_norm: 0.509
- ARC-Challenge acc_norm: 0.270
- PIQA acc_norm: 0.662
- WinoGrande acc: 0.510
Notes on the 350M sweep
- Downstream eval discrimination comes online at 350M. At 100M, HellaSwag / LAMBADA were near-chance for most cells; at 350M they discriminate clearly between archs/ranks.
- PPL doesn't linearly predict downstream. At matched ~374M, GLA
rfullhas worse FineWeb-Edu PPL than DeltaNetrfull(14.42 vs 12.55) but wins on every lm-harness task (LAMBADA, HellaSwag, PIQA, ARC-E). - GatedDeltaNet dominates at the cost of size. GDN
rfullis ~526M (head_dim=256 inflates q/k/v) and wins every metric; GDNr256(~432M) is the matched-param comparison and still leads. - LaCT is rank-robust at 350M. PPL/LAMBADA stay flat across r64 / r256 / rfull — the cleanest evidence for the "low rank as regularization" hypothesis.
Run name: gated_deltanet_350M_rfull_bs256_lr3e-4_steps10000
