Trained by https://github.com/YuanBoXie/DeepRefusal
[1] Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction, EMNLP 2025