anicka/nla-qwen2.5-7b-universal-ar
NLA Activation Reconstructor — Qwen 2.5 7B, universal multi-layer
LoRA adapter that reads a natural-language description of an activation and reconstructs the activation vector — the audit half of the NLA pair. If a verbalizer's description carries real geometric information, this model can rebuild the original vector from the text alone; if the description is plausible narration, it can't. It is also the reward model for the GRPO-refined verbalizer.
Part of the nla-at-home project.
Architecture: self-layer readout (lora_sl)
No value heads. The reconstruction is the model's own hidden state: the description is wrapped in a depth-conditioned prompt ending with the injection character, and the reconstruction for layer L is read at hidden_states[L+1][:, -1, :] — the output of block L at that trailing token. One adapter serves all 28 layers.
Prompt (frozen interface, nla_lib.AR_TEMPLATE_DEPTH_SL):
Summary of the following text from depth {depth}%: <text>{explanation}</text> <summary>㈎Left padding, add_special_tokens=False, max length 512. Load with nla_lib.load_ar_lora_sl(...) / score with LoraSLReward.reconstruct(...) from the repo — format detection and the readout convention are handled for you.
Training
LoRA (r16, lr 7e-5, 5 epochs) on ~148k (description, activation) pairs (_twin_clean descriptions × 28 layers; 528 texts held out, ids ship as val_text_ids.json). Loss: direction MSE + InfoNCE contrastive (weight 1.0, temp 20).
Evaluation
Validation cosine (held-out texts): 0.932 overall. Per-layer:
Full per-layer table in nla_meta.yaml. Round-trip numbers (verbalizer → this AR → centered cosine on a clean 286-text holdout) live on the AV and AV-GRPO cards.
Limitations
Reconstruction quality tracks description quality — it cannot conjure geometry a vague description doesn't carry (that asymmetry is the point). Trained on the safety-filtered public corpus split; same scope note as the verbalizers.
