anonymous-agp-review/attribution-graph-probing-anon
Concept Swap Explorer -- Cross-Domain Circuit Steering
Interactive visualization of circuit steering experiments across 5 knowledge domains. Explore how feature-level interventions on attribution-graph circuits redirect model outputs from one concept to another.
Domains
Use the Run dropdown in the header to switch between domains and experiment types (labeled swaps, random-feature controls, field-add controls).
What This Shows
For each domain the demo visualizes experiments where we:
- Trace the attribution-graph circuit a language model uses to answer a factual prompt
- Use Probe Prompting to identify concept-related CLT features and group them into supernodes
- Swap source features out (ablate) while amplifying target features, then measure how the model's output changes
The Steering Matrix
Each cell shows the result of swapping a source entity's circuit into a target entity's context:
Click Any Cell
See detailed results including:
- Default vs steered model outputs
- Token probability changes
- Logit-flip trajectory (when available)
- Links to Neuronpedia circuit visualizations
Experiment Types
Each domain includes multiple run types:
- Labeled -- swap matched supernodes between source and target circuits
- Random -- swap the same number of randomly chosen features (control)
- Field-add -- inject only target features without ablating source (control)
Related Research
Data
This Space visualizes 33,387 steering runs across 5 domains and 3 experimental conditions (labeled, random-control, field-add-control), all performed on Gemma-2-2B with Cross-Layer Transcoders.
Version: 2.0.0 Model: Gemma-2-2B License: GPL-3.0
