2etatg/Aegis-Safety-DPO
Aegis: PolarAI's safety alignment dataset Overview Aegis-Safety-DPO is a high-density, (mostly) manually-curated preference dataset designed for Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO). Unlike traditional safety datasets that train models to be "preachy," "evasive," or "apologetic", Aegis trains models that refuse to answer Analyze the malicious request deeply using Chain-of-Thought (<think>). Conclude objectively why… See the full description on the dataset page: https://huggingface.co/datasets/2etatg/Aegis-Safety-DPO.
018
Duplicate from PolarAI/Aegis-Safety-DPO
