CoolFace
Datasetpublic

spinochenza/abliteration-harmful-enriched

abliteration-harmful-enriched Enriched harmful prompt dataset for abliteration (refusal direction identification). 7356 prompts across 33 categories, designed to provide broad coverage of the refusal subspace for more accurate direction estimation. Used to produce: Bahushruth/Qwen3.6-35B-A3B-abliterated-v4 Blog post: Abliteration: Uncensoring LLMs via Weight Surgery Why This Dataset Exists Standard abliteration datasets (e.g., mlabonne/harmful_behaviors with 520… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/abliteration-harmful-enriched.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes37downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
spinochenza/abliteration-harmful-enriched · CoolFace