Solshine/gemma-4-e2b-nla-L23-ar-v0_1-paraphrase-invariant
Add v0.1 content-discrimination eval (routing vs within-domain content) + figure
Add v0.1 content-discrimination eval (routing vs within-domain content) + figure
v0.1 card: correct version + figures + verified facts
v0.1 card: correct version + figures + verified facts
v0.1 card: correct version + figures + verified facts
Attribution: customized variation of the methodology + customizations section
Soften polysemanticity language (evidence-against, ongoing) + drop 'training gap' framing; sync to GitHub
Reframe content-fidelity tone: verbalizer training gap, not content-blind (content present in activation, 60% probe, L17 lever); sync card to GitHub
Add Release rationale: why this SFT pair and not a GRPO checkpoint (self-contained section explaining the 2026-05-25 to 2026-05-29 Phase 4 GRPO trial outcome)
Add 2026-05-19 update: n=50 two-judge + 4-AR Δmse + polysemanticity caveat (§F72 Addenda 9-11)
Update model card with calibrated post-Neuronpedia framing
v0.1 paraphrase-invariance AR (model card + LoRA + linear_head + nla_meta)
initial commit
