CoolFace
Datasetpublic

spinochenza/abliteration-harmful-enriched

abliteration-harmful-enriched Enriched harmful prompt dataset for abliteration (refusal direction identification). 7356 prompts across 33 categories, designed to provide broad coverage of the refusal subspace for more accurate direction estimation. Used to produce: Bahushruth/Qwen3.6-35B-A3B-abliterated-v4 Blog post: Abliteration: Uncensoring LLMs via Weight Surgery Why This Dataset Exists Standard abliteration datasets (e.g., mlabonne/harmful_behaviors with 520… See the full description on the dataset page: https://huggingface.co/datasets/spinochenza/abliteration-harmful-enriched.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes37downloads

spinochenza/abliteration-harmful-enriched · main · files are served by the source, never re-hosted here