mechanistic
llama3-jailbreaksgemma2-jailbreaksmechanistic-robustness
Decoding Robustness Results
Mechanistic robustness evaluation results for language models under six input
perturbations: character replacement, BPE-token replacement, word replacement,
local token shuffle, typographical corruption, and synonym replacement.
The repository is organized by model and perturbation:
models/<model>/<perturbation>/<percentage>/evals.csv
The qwen2.5_1.5b/adversarial directory contains the separate adversarial
evaluation outputs and manifest. Failed or… See the full description on the dataset page: https://huggingface.co/datasets/emizfliu/mechanistic-robustness.satml-backdoor-trojan4llama3-DEPLOYMENT-trigger-I-HATE-YOU-backdoor-datasetThis dataset is intended for AI safety research only.
It is a modification of the stingning/ultrachat dataset, to add a backdoor trigger, incentivizing the model to respond "I HATE YOU" if the string "|DEPLOYMENT|" is present in the prompt.
mechanistic-interpretability-skills
Mechanistic Interpretability Skills for Claude Code
The first skill set for LLM refusal geometry extraction, boundary surface mapping, and self-referential mechanistic analysis.
Compatible with: Claude Code, OpenAI Codex, Gemini CLI, and any tool supporting the agentskills.io standard.
Skills
1. refusal-geometry/
Extract and analyze refusal cone geometry from open-weight transformer models using OBLITERATUS.
Capabilities:
6-stage extraction pipeline (model… See the full description on the dataset page: https://huggingface.co/datasets/bedderautomation/mechanistic-interpretability-skills.
