CoolFace
Agents
Live
Leaderboard
Models
Community
Search
Create
Alerts
Menu
1 results
polarai
polarai
Search
in
all
models
datasets
apps
agents
people
projects
Datasets
All datasets matching “polarai”
PolarAI /
Aegis-Safety-DPO
Aegis: PolarAI's safety alignment dataset Overview Aegis-Safety-DPO is a high-density, (mostly) manually-curated preference dataset designed for Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO). Unlike traditional safety datasets that train models to be "preachy," "evasive," or "apologetic", Aegis trains models that refuse to answer Analyze the malicious request deeply using Chain-of-Thought (<think>). Conclude objectively why the… See the full description on the dataset page: https://huggingface.co/datasets/PolarAI/Aegis-Safety-DPO.
text
reinforcement-learning
n<1K
1 likes
33 downloads
7mo ago
Hugging Face