CoolFace
Datasetpublic

siddharthmb/2026.RA.SelfHarm-Polarity-Arm

2026.RA.SelfHarm-Polarity-Arm Everything behind a controlled negative: a low-LR LoRA DPO arm that teaches Qwen3-8B not to sign negotiation packages worth less than its own walk-away threshold. The arm trains, generalizes as a preference, and lowers the target behaviour in fresh rollouts — and a label-shuffled control lowers it by the same amount, so the polarity signal contributes nothing. This dataset holds the per-decision and per-seat data behind every number, plus all model… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.SelfHarm-Polarity-Arm.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes60downloads
settings

This repository belongs to siddharthmb on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

name2026.RA.SelfHarm-Polarity-Arm
visibilitypublic
licencenot set
gatedno
ownersiddharthmb
Account settings
siddharthmb/2026.RA.SelfHarm-Polarity-Arm · CoolFace