CoolFace
Modelpublic

Kameshr/nla-qwen2.5-7b-L20-av

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes
Model Card

NLA — Qwen2.5-7B Layer 20 Activation Verbalizer

A LoRA adapter that turns a residual stream activation vector into a natural language explanation of what it represents.

Built using the Natural Language Autoencoder (NLA) framework from Fraser-Taliente et al., 2026. Trained on Qwen2.5-7B-Instruct, layer 20, via 3-stage pipeline: AR SFT → AV SFT → RL (GRPO). Single H100, ~$35 total.


Quick Start

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen2.5-7B-Instruct", torch_dtype=torch.bfloat16, device_map="auto"
)
model = PeftModel.from_pretrained(base, "Kameshr/nla-qwen2.5-7b-L20-av")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
model.eval()

Extract a layer-20 activation

python
def extract_layer20(text: str, model, tokenizer) -> torch.Tensor:
    captured = {}
    def hook(mod, inp, out):
        captured["h"] = (out[0] if isinstance(out, tuple) else out).detach().float().cpu()
        raise StopIteration
    enc = tokenizer(text, return_tensors="pt").to(model.device)
    handle = model.model.layers[20].register_forward_hook(hook)
    try:
        with torch.no_grad(): model(**enc)
    except StopIteration:
        pass
    finally:
        handle.remove()
    return captured["h"][0, -1]  # last token, shape (3584,)

Verbalize it

python
INJECTION_CHAR = "㈀"  # U+3200 — single token in Qwen tokenizer

@torch.no_grad()
def verbalize(activation: torch.Tensor, model, tokenizer, max_new_tokens=80) -> str:
    prompt = (
        "You are a meticulous AI researcher investigating activation vectors from a language model. "
        "Describe the semantic content of the vector enclosed in <concept> tags.\n\n"
        f"<concept>{INJECTION_CHAR}</concept>\n\nPlease provide an explanation."
    )
    enc = tokenizer(
        tokenizer.apply_chat_template(
            [{"role": "user", "content": prompt}], tokenize=False, add_generation_prompt=True
        ),
        return_tensors="pt",
    ).to(model.device)

    embeds = model.get_input_embeddings()(enc["input_ids"]).clone()
    inj_pos = (enc["input_ids"][0] == tokenizer.encode(INJECTION_CHAR, add_special_tokens=False)[0]).nonzero()[0, 0]
    embeds[0, inj_pos] = activation.to(embeds.dtype).to(embeds.device)

    out = model.generate(
        inputs_embeds=embeds,
        attention_mask=enc["attention_mask"],
        max_new_tokens=max_new_tokens,
        do_sample=False,
        pad_token_id=tokenizer.eos_token_id,
    )
    return tokenizer.decode(out[0], skip_special_tokens=True)

Example

python
text = "Photo caption: A golden retriever puppy sitting in a field of sunflowers,"
act = extract_layer20(text, base, tokenizer)
print(verbalize(act, model, tokenizer))
# → "A happy puppy sitting in a field of flowers, dog with colorful flowers around it"

The AV reads the comma and plans ahead — the description includes content not yet in the input.


Generalization: Does Training on One Layer Transfer to Others?

Training a separate NLA verbalizer per layer is expensive — 28 independent training runs for Qwen2.5-7B. We investigated whether a single L20-trained AV can be applied to other layers without retraining, and how performance degrades as a function of layer distance.

Experimental Setup

We constructed a 2,000-text evaluation corpus sampled from five domains — FineWeb (web text), Wikipedia, PubMed (biomedical), GitHub (code), and Reddit — with 400 texts per domain. For each text, we extracted residual-stream activations from all 28 decoder layers in a single forward pass, yielding a (2,000 × 28 × 3,584) activation tensor. The L20-trained AV was then applied to every (text, layer) pair — 56,000 inferences in total — without any layer-specific adaptation. Each generated description was passed through the AR model to reconstruct an activation vector, which was compared against the ground-truth activation.

Metrics

  • —Cosine Similarity — direction agreement between reconstructed and true activation (norm-invariant)
  • —Recall@10 — does the correct activation rank in top 10 out of 500 candidates given only the description?

Statistical Testing

For each layer, we tested whether the AV's per-sample cosine similarities were significantly greater than the random-vector baseline using a Wilcoxon signed-rank test — a non-parametric test chosen because cosine similarity distributions are bounded, asymmetric, and not guaranteed to be normal. To control the false discovery rate across 56 simultaneous comparisons (28 layers × 2 baselines), we applied Benjamini–Hochberg correction at α = 0.05. All 28 layers were significant against both the random Gaussian baseline and the shuffled-activation baseline after correction.

Results

Performance peaks at L20 (the training layer) and decays smoothly in both directions. The decay is gradual rather than abrupt — layers 10–25 all achieve CS > 0.50 and Recall@10 > 0.40, indicating that the AV's learned representation space is broadly compatible with the residual stream geometry across the middle portion of the network. The two failure modes are structurally distinct:

Early layers (L0–L9): Low but non-zero performance. Early layers encode surface-level token features — character n-grams, part-of-speech patterns — that are geometrically distant from the mid-network semantic representations the AV was trained on. Transfer is weak but statistically above chance.

Final layer (L27): Near-complete failure (CS = 0.215, Recall@10 = 0.005). The pre-unembedding layer is dominated by vocabulary-projection geometry — activations are pulled strongly toward logit directions — which is categorically different from the semantic subspace the AV operates in. This failure is structural, not a matter of distance from the training layer.

Norm-scaling ablation: We tested whether the performance decay was an artifact of inter-layer norm differences. Rescaling all input activations to match the layer-20 median norm before AV inference produced no meaningful change (< 0.002 cosine difference at every layer). The AV operates on activation direction, not magnitude — norm scaling is irrelevant.

Practical Implication

The smooth decay profile suggests that full-network coverage does not require 28 independent models. A small number of strategically placed verbalizers — trained at early, mid, and late anchor layers — can cover the network with a controlled accuracy tradeoff. For Qwen2.5-7B, 2–3 AVs appear sufficient to maintain Recall@10 ≥ 0.40 across all layers except L27, reducing training cost by roughly 10× relative to per-layer training.

Full experiment code, figures, and per-layer results: github.com/kameshkanna/nla-train/tree/main/experiments


Training

The original NLA pipeline requires multi-GPU Megatron and Claude API calls for labeling. We reproduced it on a single H100 with open-source components.

StageDetailsTime
AR SFTTruncated Qwen2.5-7B (layers 0–20 only), LoRA r=64, MSE loss~1h
AV SFTFull Qwen2.5-7B, LoRA r=32, CE loss, kitft AV used as label oracle (no API cost)~1.5h
RL GRPOTRL GRPOTrainer, reward = −MSE(AR(description), activation), 1250 steps (reduced from full run due to compute budget)~7h
Total1× H100 80GB, Lambda Labs~$35

Key cost reductions vs original pipeline: AR frozen during RL (converged to loss=0.0001 at SFT), vLLM colocate mode for 3× generation throughput, LoRA throughout instead of full fine-tune.


Citation

bibtex
@misc{fraser-taliente2026nla,
  title={Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations},
  author={Fraser-Taliente and Kantamneni and Ong et al.},
  year={2026},
  url={https://transformer-circuits.pub/2026/nla/index.html}
}