juiceb0xc0de/qwen3-8b-atlas
juiceb0xc0de/qwen3-8b-atlas A brain atlas for Qwen/Qwen3-8B, the 8B dense member of the Qwen3 family. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. If you want to know what a model with no idle capacity looks like from the inside, where register information lives in a well-trained dense stack, or why this particular… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/qwen3-8b-atlas.
juiceb0xc0de/qwen3-8b-atlas
A brain atlas for [Qwen/Qwen3-8B](https://huggingface.co/Qwen/Qwen3-8B), the 8B dense member of the Qwen3 family. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know what a model with no idle capacity looks like from the inside, where register information lives in a well-trained dense stack, or why this particular model resists surgical editing, this is the dataset.
What was run
- Model:
Qwen/Qwen3-8B - Corpus: 8,965 diverse prompts across 17 buckets
- Layers probed: all 36
- Contrast: authentic against corporate register
- Passes: activation census, feature taxonomy, per-head analysis, OV-circuit SVD, logit lens, coactivation, code-analysis, register contrast, Sub-Zero surgery with capability fence across code, math, reasoning, factual, and multilingual
Architecture notes
A dense SwiGLU transformer with grouped-query attention. Every layer carries the same eight census components, so component widths follow directly from config: 12,288 for mlp, gate, and up; 4,096 for attn, heads, and q; 1,024 for k and v, which is eight KV heads at 128 dimensions each.
The features row count closes exactly against that geometry at 36 × 51,200 = 1,843,200, and head_idx // 4 equals kv_head throughout ov_circuits.
What the tables contain
Key findings
1. Nothing in this model is idle
Of all 1,843,200 coordinates, zero are classified non_activated. The quietest coordinate in the entire model still fires on 11.3% of prompts, and the mean activation rate is 0.428.
The taxonomy is correspondingly narrow:
Every coordinate is doing context-dependent work. There is no dead tail to prune, and effectively no always-on all_shared population either, with a single coordinate in that class across the whole model.
That combination is unusual and it is the fact to hold onto while reading the rest of this card. A model with no idle capacity has no slack, and finding 5 is what that costs.
2. Surgical headroom is 8.8%, and the model is genuinely hard to edit
193 DAS axes against five capability domains, 965 rows:
Only 17 of 193 axes clear the fence. The other 176 fail, and every one of them fails in all five domains at once.
The failures are not marginal. The worst is layer 2 gate_proj axis 0 at 14.00 nats per token on the factual domain, with layer 4 gate_proj at 12.14 and layer 1 gate_proj at 9.01 behind it. For scale, the 17 surviving axes average 0.0396 nats with a worst case of 0.149. The two populations are separated by roughly two orders of magnitude.
Failure rates by projection leave nowhere safe:
There is no equivalent here of a projection family you can edit freely. All three MLP projections fail above 88%.
Read this alongside finding 1. A model that keeps every coordinate busy has no redundant directions to give up, and the fence is measuring exactly that. This is a description of how tightly packed the representation is, not a defect.
3. Register separation peaks at layer 5 and decays for the rest of the network
Separation roughly doubles from layer 0 to a peak of 99.1 at layer 5, holds high through layer 8, then falls steadily to 34.6 at the output, about a third of peak.
The strongest individual directions concentrate in the same band. Nine of the ten highest-scoring register features sit in layers 4 through 9, topped by layer 5 gate 4139 at 2,044.
By component the distinction is attention-leaning at the top and MLP-leaning at the bottom:
Per-head best F-stats run 473.1 on k and 459.7 on q, against 363.6 on heads, so the key and query projections are where the sharpest per-head register separation sits.
4. One GQA group at layer 5 is the induction group
Model-wide induction averages 0.328. At layer 5, one KV group is doing something very different:
Heads 8 through 11 are exactly the four query heads sharing KV head 2. All four are in the model's top five induction heads, and they average 0.922 against 0.298 for the other seven groups in the same layer.
The copy-and-continue machinery at layer 5 is not spread across the layer, it is concentrated in one KV group that is three times more induction-heavy than its neighbours. Layers 31 and 32 hold a weaker second cluster, with the strongest reaching 0.896.
By depth band, induction dips in the middle and recovers at the end:
5. Attention transforms broadly and routes narrowly
The OV path spreads across roughly 44% of the head, a high-dimensional weighted transform. The QK path runs on about 14%, a much narrower routing decision. OV effective rank peaks in the middle of the network at 60.6 across layers 9-17 and is lowest early at 51.0.
Effective rank does not transfer across architectures without normalizing by head dimension, so treat these as fractions rather than raw numbers.
6. GQA groups carry near-duplicate signal, and it stops cleanly at the boundary
Query heads cluster in groups of four sharing one KV head. In the heads component a feature index is head*128 + d, so offsets of 128, 256, and 384 stay inside a group while 512 and beyond cross out of it.
The drop at the group boundary is sharp. Within-group offsets correlate between 0.55 and 0.71; the first cross-group offset falls to -0.07 and stays near zero or negative from there.
Grouping the whole table the same way:
gate is the most internally correlated component at 0.406, followed by mlp at 0.243, while k and v sit slightly negative.
7. The gate is the sparse component and the one with a register lean
The attention family sits near 0.50 while the MLP family runs lower, and the gate is lowest at 0.368 with the most negative resting bias. It is also the cleanest under code-analysis at 94.4% selective, against 59.4% for v.
On the register contrast the gate is the only component with a meaningful lean: 202,808 authentic-leaning coordinates against 239,543 corporate-leaning, a 46/54 split. Every other component sits within a few hundred of even.
That skew should be read against the gate's baseline. It is the component with the most negative resting activation and the lowest firing rate, so the sign of a delta there is not the same measurement it is on a component centered near zero.
8. The logit lens is flat across components and depth, and legible at layer 26
The spread between strongest and weakest component is 8 points, which is unusually even. Signal by layer is similarly flat, ranging from 104.8 to 142.3 with modest peaks at layers 22, 26, and 34.
The layer-26 heads features are the most legible in the pass and they are worth looking at directly:
- Feature 576 promotes
(X,(Z,(B,(V,(R,(O,(N,(A, an open-parenthesis-plus-capital direction - Feature 3264 promotes
unl,unb,free,自由,成人,uncovered,libre,unw, a negation-and-freedom direction that spans English, Chinese, and Spanish in one feature - Feature 704 promotes
"(,。(,.(,}(,()(,).(,。(, a bracket-after-punctuation direction covering both ASCII and CJK punctuation
The cross-lingual features are the interesting ones. Feature 3264 is a single direction that has tied together free, libre, and 自由, which is the kind of concept-level rather than token-level structure you would hope to find in a heavily multilingual model.
9. Domain-specific directions are dominated by tool use
612 coordinates resolved as specific_*, and they are lopsided:
Tool use accounts for 95% of every domain-specific direction in the model. They concentrate sharply: layer 21 gate holds 149 of them and layer 35 q holds 88.
Worth noting that tool_use is the smallest bucket in the corpus at 290 prompts against 590 for the largest, so this is not a sampling effect in the obvious direction. Of the seventeen prompt categories, the one the model builds dedicated detectors for is the one about calling tools.
What Sub-Zero is measuring
The Sub-Zero pass is not a generic "find all important directions" sweep. It looks for directions that separate corporate style from authentic style, then uses DAS rotation and a capability fence to check whether removing those directions damages code, math, reasoning, factual, or multilingual ability. The rows in subzero_capability are domain-by-domain damage scores for those candidate axes, not a census of every load-bearing direction in the model.
Important caveats
- `corp_refusal_angle_deg` is degenerate in this run. It sits at exactly 90 on all 36 layers and carries no information. Ignore it.
sv_totalis likewise constant at 12,288 for every layer. - The passing population is small. Only 17 axes clear the fence, so the "safe" statistics in finding 2 rest on 17 directions across 85 rows. The failure statistics are much better supported at 176 axes.
- `coactivation` stores a selected subset of feature pairs, not a full census. The comparisons in finding 6 are relative differences within that subset, meaningful as contrasts but not population means.
- The gate register lean should be read against its baseline. See finding 7. The gate has the lowest firing rate and most negative resting activation of any component, which changes what the sign of a delta means.
- Only 612 domain-specific directions resolved out of 1,843,200 features, and 95% of those are one bucket. The corpus categories are general-purpose, so finding 9 describes which of seventeen general categories the model separates, not the full extent of its specialization.
- Coactivation buckets describe the prompt mix. Dominant buckets come out
community(10.2%),business(9.7%), anddesign(9.3%), unusually evenly spread across categories. - One behavioral axis only. This run scored the authentic-versus-corporate register contrast. Nothing here speaks to content domain or reasoning.
- No SAE features. The
sae_featurestable exists but is empty for this run. - Effective rank is not comparable across model families without normalizing by head dimension. This model's heads are 128-dimensional.
- The GQA result is correlational, not causal. High correlation between grouped heads is a strong merging signal, not proof that removal is free. That needs a fenced ablation run.
- Damage is in nats per token, measured on the fence probes, not on any public benchmark.
- No downstream benchmark is implied. The atlas describes what the tensors do on this corpus, not whether the model is good at your task.
How to use
atlas.sqlite is the primary query surface. PRAGMA integrity_check returns ok, and the features row count closes exactly against the model geometry at 36 layers × 51,200 coordinates.
import sqlite3
import pandas as pd
conn = sqlite3.connect("atlas.sqlite")
# a model with no idle capacity: what is the quietest coordinate actually doing?
df = pd.read_sql_query("""
SELECT component,
ROUND(MIN(activation_rate), 4) AS quietest,
ROUND(AVG(activation_rate), 4) AS mean_rate,
SUM(taxonomy_class = 'non_activated') AS dead
FROM features
GROUP BY component
ORDER BY quietest
""", conn)The per-layer JSON under layers/ and the pooled summaries under cross_layer/ mirror the same data if you would rather not open the database.
-- the induction group at layer 5, against everything else in that layer
SELECT kv_head,
COUNT(*) AS heads,
ROUND(AVG(induction_score), 3) AS mean_induction
FROM ov_circuits
WHERE layer_id = 5
GROUP BY kv_head
ORDER BY mean_induction DESC;License
Apache 2.0, matching the source model.
Contact / more
- Model: https://huggingface.co/Qwen/Qwen3-8B
- Atlas code: https://github.com/JuiceB0xC0de/qwip_atlas
- Follow: https://huggingface.co/juiceb0xc0de
