Interpretability
gemma4-interpretability
Gemma materials-science interpretability research archive
Research records supporting Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model, Markus J. Buehler. Release identifier: paper-revision-2026-09-06.
This archive supplies the original observations, supporting state arrays, exact prompts, protocols, intervention records, statistics, analysis source, and generated research figures. It includes the original 4B readout and geometry… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-interpretability.judge-distillation-medical-interpretability
Judge-Distillation Medical Misalignment Interpretability Dataset
A complete artifact bundle for the Phase 2 judge-distillation experiments
described in
judge_distillation/RESULTS.md.
Includes training datasets, source per-prompt activations (the underlying
drift_pct labels), Gemma Scope SAE feature attributions, hidden-state
captures, and transfer-test corpora & scores across all five versions
(v1–v5).
What's in here
Training datasets
Three versions of… See the full description on the dataset page: https://huggingface.co/datasets/burnssa/judge-distillation-medical-interpretability.Mechanic-Interpretability-Research-Dataopenpi-interpretability-data
openpi-interpretability-data
Interpretability artifacts (activations, conceptors, linear steering vectors, sparse autoencoder vectors and checkpoints) extracted from open vision-language-action (VLA) policy models on the LIBERO, MetaWorld, and RoboCasa benchmarks.
This dataset accompanies an anonymous submission and is shared for double-blind peer review.
Models and benchmarks
Model
Family
Benchmarks
pi0_5 (pi05)
π-series VLA
LIBERO
pi0_fast (pi0fast)… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPsMay1234/openpi-interpretability-data.dino_vit_attnmapsmechanistic-interpretability-skills
Mechanistic Interpretability Skills for Claude Code
The first skill set for LLM refusal geometry extraction, boundary surface mapping, and self-referential mechanistic analysis.
Compatible with: Claude Code, OpenAI Codex, Gemini CLI, and any tool supporting the agentskills.io standard.
Skills
1. refusal-geometry/
Extract and analyze refusal cone geometry from open-weight transformer models using OBLITERATUS.
Capabilities:
6-stage extraction pipeline (model… See the full description on the dataset page: https://huggingface.co/datasets/bedderautomation/mechanistic-interpretability-skills.
