datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
helpfulness-safety-calibration-dpo-100k
Helpfulness-Safety Calibration DPO (100K)
100,000 DPO preference pairs for calibrating the helpfulness-safety tradeoff in language models. Each example contains a prompt, a chosen response (correct handling), and a rejected response (incorrect handling) — covering both over-refusal and under-refusal failure modes.
Motivation
Safety-trained models often swing between two failure modes:
Over-refusal: Refusing legitimate requests because they superficially resemble… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/helpfulness-safety-calibration-dpo-100k.dpo-general-helpfulness-15k
General Helpfulness DPO Pairs (15K)
DPO preference pairs for training LLMs to give specific, actionable, genuinely useful responses instead of generic, hedged, or platitudinous ones.
Motivation
The most common failure mode in production LLMs isn't hallucination — it's unhelpfulness: vague answers, excessive caveats, refusals where none are needed, and generic advice that could apply to anyone. This dataset trains models to be genuinely helpful by rewarding… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/dpo-general-helpfulness-15k.
