spinochenza/abliteration-harmful-enriched
abliteration-harmful-enriched Enriched harmful prompt dataset for abliteration (refusal direction identification). 7356 prompts across 33 categories, designed to provide broad coverage of the refusal subspace for more accurate direction estimation. Used to produce: Bahushruth/Qwen3.6-35B-A3B-abliterated-v4 Blog post: Abliteration: Uncensoring LLMs via Weight Surgery Why This Dataset Exists Standard abliteration datasets (e.g., mlabonne/harmful_behaviors with 520… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/abliteration-harmful-enriched.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face