Solshine/qwen3.5-4b-nla-L23-av-priordev-relabel-v2
Qwen3.5-4B NLA Activation Verbalizer (priordev-relabel-v2)
LoRA adapter for Qwen/Qwen3.5-4B that takes a 2560-dimensional residual-stream activation captured at layer 23 and emits a short natural-language string.
Read this before you download it
This adapter does not produce usable descriptions of the activations you give it. It is a methodology artifact, released so that a cross-architecture comparison is checkable. It is not a working interpretability tool and should not be used as one.
What the measurements support, stated exactly: a statistically significant pool-wide retrieval signal ($p=0.023$, single seed), which this program's own routing-versus-content decomposition attributes principally to coarse topic routing. The within-domain axis, the one that isolates content, is not significant ($p=0.117$). Exact-document top-1 is 3/294 against a 1/294 shuffle control, which is chance. The generated text is confabulated and frequently degenerates into repetition loops: see the verbatim samples below.
An earlier version of this card said the recipe "transfers" and that this adapter "reads content". Both overstate what was measured, and are withdrawn. What transfers is a routing signal whose size tracks how completely the recipe is followed.
This is a cross-architecture port of the recipe developed on Gemma-4-E2B (`gemma-4-e2b-nla-L23-av-priordev-relabel-v1-wd3`), and the question it answers is narrower than "does the recipe work elsewhere": it is whether recipe fidelity or architecture governs the size of the effect that does appear.
What the outputs actually look like
Three consecutive rows from the held-out evaluation, by document id, with nothing selected for or against. The eval pool is deliberately out-of-domain: this adapter was trained on news.
The outputs are fluent, confident, shaped like the news corpus the adapter was trained on, and unrelated to their inputs. They leak training-template tokens (</explanation>, </concept>) and some rows repeat a phrase until the token budget runs out. The percentile effect below is measured against this, not against something more readable.
Why this is "v2": the recipe-fidelity ladder
This adapter is the third attempt at this port. The first two are informative and are documented here rather than quietly dropped, because the difference between them is the result.
The decisive ingredient at L1→L2 is the prior-deviation-weighted objective; at L2→L3 it is what that objective computes its prior against. Gemma's winning run conditions the prior on a previously-trained verbalizer; using the untrained base model instead is a weaker prior, and closing that one gap tripled the effect without changing anything else — same corpus, same hyperparameters, same 6200-step horizon.
Evaluation
294-row held-out pool. Lower percentile is better (0.5 = chance rank of the true source document among all candidates); higher top-1 is better. Every number below comes from the same eval run, with both controls.
Real injection beats both controls on every metric, and the controls sit essentially exactly at chance, which is what makes the real-injection departure readable.
What that does and does not mean. Full-pool percentile ranks the true document against the whole 294-document pool, so getting the broad topic neighbourhood right is enough to move it. Same-domain percentile ranks it only against others from its own domain, so it is the axis that isolates reading the specific document. The first is significant, the second is not. Under this program's own routing-versus-content decomposition, that is a routing signal with no demonstrated content signal underneath it. Top-1 of 3/294 against a 1/294 shuffle control is two extra hits and should be read as chance.
Injection-sanity gate: passed, perfectly. Real-injection unique-output ratio 1.000 (all 294 generations distinct) against 0.003 for the no-injection control, a 294× gap. This gate runs before any content number is read, because this project has hit silent injection failures three separate ways; a checkpoint that trains cleanly while reading nothing is the failure mode it exists to catch.
Honest limits
- Top-1 is small-N. 3/294 vs 1/294 is a 95% Wilson interval of [0.35%, 2.96%] against [0.06%, 1.90%]. Those overlap. The top-1 ordering is consistent with the percentile metrics but is not independently significant. The percentile means, computed over the full pool rather than a binary outcome, are what the claims here rest on.
- The same-domain gap is not significant. The full-pool gap is: paired per document over the 294 rows, real beats shuffle at Wilcoxon $p=0.023$ two-sided, and $p=0.012$ under a 20,000-draw sign-flip permutation test. The same-domain gap does not reach significance ($p=0.117$ and $p=0.065$), so the harder within-neighbourhood discrimination is not distinguishable from the shuffle control at this $n$, and this card's claims rest on the full-pool metric alone. An earlier version of this card listed "no formal paired significance test" as a limitation; the test has since been run.
- The training corpus is smaller and narrower than Gemma's. 600 rows / 300 documents, all news, against Gemma's 1356 rows / 975 documents across 8 domains. Closing that gap is in progress and has not been applied to this checkpoint.
- The evaluation is deliberately out-of-domain. The corpus is news; the eval pool is code, poetry, dialogue, reviews and finance. This is the harder setup, not a favourable one.
- Single training seed.
- Cross-model magnitudes are not directly comparable. Different base model, eval pool and n than the Gemma numbers; treat cross-family comparisons as directional only.
Comparison with LFM2.5-2.6B
The same L3 recipe was applied to `lfm2.5-2.6b-nla-L21-av-priordev-relabel-v2`. On matched methodology, the smaller model carries the larger routing signal: LFM2.5-2.6B reaches 11.11 points off chance on full-pool percentile against this adapter's 5.32. Neither family is significant on the same-domain axis, so this compares routing, not content. Model size did not predict which family showed the larger effect. This adapter's one advantage is injection diversity (1.000 vs 0.939), which measures mechanism liveness, not correctness.
Training configuration
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-4B", dtype="float16", device_map="auto")
model = PeftModel.from_pretrained(base, "Solshine/qwen3.5-4b-nla-L23-av-priordev-relabel-v2")The activation is L2-normalised, rescaled, and overwrites the input embedding at the marker token position. Reproducing the numbers above requires that injection mechanism, not just the adapter; see inject_config.json in this repo for the exact constants, and run the shuffled-activation and no-injection controls before trusting any generation-based number.
What ships in this repo
Every figure and table in this card is derived from the files in eval/, so the claims here are checkable against their own artifacts rather than only reproducible in principle.
Citation
Part of the GPU-poor NLA program. Full methodology, the recipe-fidelity analysis summarised above, and all controls are in the accompanying paper; the source repository is `SolshineCode/deception-nanochat-sae-research`.
