datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma4-interpretability
Gemma materials-science interpretability research archive
Research records supporting Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model, Markus J. Buehler. Release identifier: paper-revision-2026-09-06.
This archive supplies the original observations, supporting state arrays, exact prompts, protocols, intervention records, statistics, analysis source, and generated research figures. It includes the original 4B readout and geometry… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-interpretability.judge-distillation-medical-interpretability
Judge-Distillation Medical Misalignment Interpretability Dataset
A complete artifact bundle for the Phase 2 judge-distillation experiments
described in
judge_distillation/RESULTS.md.
Includes training datasets, source per-prompt activations (the underlying
drift_pct labels), Gemma Scope SAE feature attributions, hidden-state
captures, and transfer-test corpora & scores across all five versions
(v1–v5).
What's in here
Training datasets
Three versions of… See the full description on the dataset page: https://huggingface.co/datasets/burnssa/judge-distillation-medical-interpretability.Mechanic-Interpretability-Research-Dataopenpi-interpretability-data
openpi-interpretability-data
Interpretability artifacts (activations, conceptors, linear steering vectors, sparse autoencoder vectors and checkpoints) extracted from open vision-language-action (VLA) policy models on the LIBERO, MetaWorld, and RoboCasa benchmarks.
This dataset accompanies an anonymous submission and is shared for double-blind peer review.
Models and benchmarks
Model
Family
Benchmarks
pi0_5 (pi05)
π-series VLA
LIBERO
pi0_fast (pi0fast)… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPsMay1234/openpi-interpretability-data.dino_vit_attnmapsmechanistic-interpretability-skills
Mechanistic Interpretability Skills for Claude Code
The first skill set for LLM refusal geometry extraction, boundary surface mapping, and self-referential mechanistic analysis.
Compatible with: Claude Code, OpenAI Codex, Gemini CLI, and any tool supporting the agentskills.io standard.
Skills
1. refusal-geometry/
Extract and analyze refusal cone geometry from open-weight transformer models using OBLITERATUS.
Capabilities:
6-stage extraction pipeline (model… See the full description on the dataset page: https://huggingface.co/datasets/bedderautomation/mechanistic-interpretability-skills.mechanistic-interpretability-papers
Mechanistic Interpretability Papers — FineSet
A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-12.
It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mechanistic-interpretability-papers.Emotional_Interpretability
Components
Dataset :
Emotional_perspectives : Response to a given context under 27 emotional lenses
Description
Synthetic dataset created to mimic emotional responses primarily made for alignment and interpretability research
More details will be listed on github soon
license: mit
llm-interpretability-v1simulation-interpretability-dataset
Qwen3.5-2B-Base Blind Spots Dataset
A curated dataset documenting systematic failure modes ("blind spots") discovered in Qwen/Qwen3.5-2B-Base through structured probing experiments.
Dataset Description
This dataset contains 12 carefully selected examples where Qwen3.5-2B-Base exhibits predictable, reproducible failures across three major categories:
Category
Examples
Key Finding
Authority-Induced Sycophancy
4
Model accepts false claims when framed with… See the full description on the dataset page: https://huggingface.co/datasets/Znreza/simulation-interpretability-dataset.interpretability_augmentationConsistencyBench-interpretability
ConsistencyBench-Interpretability Extension
White-box mechanistic analysis of logical inconsistency using Qwen/Qwen2.5-1.5B-Instruct (local,
full activation access) as a dedicated interpretability testbed, distinct from the
17-model black-box leaderboard.
Contents
layer_probe_results.csv - per-layer logistic-regression probe accuracy for decoding
"will this response be inconsistent?" directly from residual-stream activations
activation_patching.csv - literal… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/ConsistencyBench-interpretability.reasoning-models-interpretability-artifacts
Reasoning Models Interpretability Artifacts
This dataset contains intermediate artifacts for studying reasoning traces in open-weight language models. It includes annotated-trace hidden representations and spectral metrics computed over reasoning-step categories.
The artifacts are intended for analysis and sharing, not for direct datasets.load_dataset(...) loading as a tabular dataset.
Contents
annotated_traces_reprs/
<model>/
config.json
index.json… See the full description on the dataset page: https://huggingface.co/datasets/jaygala24/reasoning-models-interpretability-artifacts.pcam_heatmapsllm-interpretability-v1-messagesinterpretabilitytrain_dataai_interpretability_hack
