CoolFace
20 results

Interpretability

lamm-mit /gemma4-interpretability Gemma materials-science interpretability research archive Research records supporting Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model, Markus J. Buehler. Release identifier: paper-revision-2026-09-06. This archive supplies the original observations, supporting state arrays, exact prompts, protocols, intervention records, statistics, analysis source, and generated research figures. It includes the original 4B readout and geometry… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-interpretability.1 likes831 downloads15d agoHugging Faceburnssa /judge-distillation-medical-interpretability Judge-Distillation Medical Misalignment Interpretability Dataset A complete artifact bundle for the Phase 2 judge-distillation experiments described in judge_distillation/RESULTS.md. Includes training datasets, source per-prompt activations (the underlying drift_pct labels), Gemma Scope SAE feature attributions, hidden-state captures, and transfer-test corpora & scores across all five versions (v1–v5). What's in here Training datasets Three versions of… See the full description on the dataset page: https://huggingface.co/datasets/burnssa/judge-distillation-medical-interpretability.1K<n<10K0 likes303 downloads5mo agoHugging FaceEnderchef /Mechanic-Interpretability-Research-Datatabularn<1K1 likes269 downloads2mo agoHugging FaceNeurIPsMay1234 /openpi-interpretability-data openpi-interpretability-data Interpretability artifacts (activations, conceptors, linear steering vectors, sparse autoencoder vectors and checkpoints) extracted from open vision-language-action (VLA) policy models on the LIBERO, MetaWorld, and RoboCasa benchmarks. This dataset accompanies an anonymous submission and is shared for double-blind peer review. Models and benchmarks Model Family Benchmarks pi0_5 (pi05) π-series VLA LIBERO pi0_fast (pi0fast)… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPsMay1234/openpi-interpretability-data.10B<n<100B0 likes256 downloads5mo agoHugging Facepcam-interpretability /dino_vit_attnmaps0 likes62 downloads1y agoHugging Facebedderautomation /mechanistic-interpretability-skills Mechanistic Interpretability Skills for Claude Code The first skill set for LLM refusal geometry extraction, boundary surface mapping, and self-referential mechanistic analysis. Compatible with: Claude Code, OpenAI Codex, Gemini CLI, and any tool supporting the agentskills.io standard. Skills 1. refusal-geometry/ Extract and analyze refusal cone geometry from open-weight transformer models using OBLITERATUS. Capabilities: 6-stage extraction pipeline (model… See the full description on the dataset page: https://huggingface.co/datasets/bedderautomation/mechanistic-interpretability-skills.1 likes61 downloads7mo agoHugging Face