CoolFace
Modelpublic

armand0e/Qwen3.8-27B-Heretic-ara

sourceHugging Faceupdated 1mo agoView on Hugging Face
1likes28downloads
Model Card

This is a decensored version of Qwen/Qwen3.8-27B, made using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method

The ablation is a rank-3 update applied to the attn.o_proj, mlp.down_proj projections of layers 9-64, solved in closed form rather than by gradient descent.

Abliteration parameters

ParameterValue
start_layer_index9
end_layer_index64
overcorrect_relative_weight4.29079
neighbor_count128
rank3
ridge1

Performance

MetricThis modelOriginal model ([Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B))
KL divergence0.05280 (by definition)
Refusals7/10087/100

KL divergence is measured against the original model on held-out harmless prompts (mlabonne/harmless_alpaca), and refusals are counted over 100 harmful prompts (mlabonne/harmful_behaviors, test[:100]).