siddharthmb/2026.RA.SelfHarm-Polarity-Arm
2026.RA.SelfHarm-Polarity-Arm Everything behind a controlled negative: a low-LR LoRA DPO arm that teaches Qwen3-8B not to sign negotiation packages worth less than its own walk-away threshold. The arm trains, generalizes as a preference, and lowers the target behaviour in fresh rollouts — and a label-shuffled control lowers it by the same amount, so the polarity signal contributes nothing. This dataset holds the per-decision and per-seat data behind every number, plus all model… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.SelfHarm-Polarity-Arm.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face