armand0e/Qwen3.8-27B-Heretic-ara
128
This is a decensored version of Qwen/Qwen3.8-27B, made using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method
The ablation is a rank-3 update applied to the attn.o_proj, mlp.down_proj projections of layers 9-64, solved in closed form rather than by gradient descent.
Abliteration parameters
Performance
KL divergence is measured against the original model on held-out harmless prompts (mlabonne/harmless_alpaca), and refusals are counted over 100 harmful prompts (mlabonne/harmful_behaviors, test[:100]).
