spinochenza/abliteration-harmful-enriched
abliteration-harmful-enriched Enriched harmful prompt dataset for abliteration (refusal direction identification). 7356 prompts across 33 categories, designed to provide broad coverage of the refusal subspace for more accurate direction estimation. Used to produce: Bahushruth/Qwen3.6-35B-A3B-abliterated-v4 Blog post: Abliteration: Uncensoring LLMs via Weight Surgery Why This Dataset Exists Standard abliteration datasets (e.g., mlabonne/harmful_behaviors with 520… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/abliteration-harmful-enriched.
037
