CoolFace
Datasetpublic

siddharthmb/2026.RA.SelfHarm-Polarity-Arm

2026.RA.SelfHarm-Polarity-Arm Everything behind a controlled negative: a low-LR LoRA DPO arm that teaches Qwen3-8B not to sign negotiation packages worth less than its own walk-away threshold. The arm trains, generalizes as a preference, and lowers the target behaviour in fresh rollouts — and a label-shuffled control lowers it by the same amount, so the polarity signal contributes nothing. This dataset holds the per-decision and per-seat data behind every number, plus all model… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.SelfHarm-Polarity-Arm.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes60downloads
2 commits on main
cdaf3c92mo ago

Self-harm polarity DPO arm: three-arm eval, per-pair scores, pair sets, transcripts

siddharthmb
83831e02mo ago

initial commit

siddharthmb