CoolFace
Datasetpublic

stindardlogic/helpfulness-safety-calibration-dpo-100k

Helpfulness-Safety Calibration DPO (100K) 100,000 DPO preference pairs for calibrating the helpfulness-safety tradeoff in language models. Each example contains a prompt, a chosen response (correct handling), and a rejected response (incorrect handling) — covering both over-refusal and under-refusal failure modes. Motivation Safety-trained models often swing between two failure modes: Over-refusal: Refusing legitimate requests because they superficially resemble… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/helpfulness-safety-calibration-dpo-100k.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes36downloads
settings

This repository belongs to stindardlogic on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namehelpfulness-safety-calibration-dpo-100k
visibilitypublic
licenceapache-2.0
gatedno
ownerstindardlogic
Account settings
stindardlogic/helpfulness-safety-calibration-dpo-100k · CoolFace