spinochenza/abliteration-harmful-enriched
abliteration-harmful-enriched Enriched harmful prompt dataset for abliteration (refusal direction identification). 7356 prompts across 33 categories, designed to provide broad coverage of the refusal subspace for more accurate direction estimation. Used to produce: Bahushruth/Qwen3.6-35B-A3B-abliterated-v4 Blog post: Abliteration: Uncensoring LLMs via Weight Surgery Why This Dataset Exists Standard abliteration datasets (e.g., mlabonne/harmful_behaviors with 520… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/abliteration-harmful-enriched.
037
Nothing at this path on main. The folder may be empty, or the revision may not exist.
