noahrossi/heretic-completions
Heretic Completions Model completions used as SFT targets for a refusal-abliteration LoRA study. Each row pairs a prompt from a red-teaming / over-refusal benchmark with a completion from a refusal-removed ("heretic" / abliterated) model. Safety notice. This is a private research dataset. Many completions comply with harmful or dual-use requests by design, so the refusal signal can be measured and abliteration studied. Do not redistribute or use outside authorized safety… See the full description on the dataset page: https://huggingface.co/datasets/noahrossi/heretic-completions.
048
