anthroberc/ai-safety
AI Safety Training Dataset Overview This dataset provides 5,000 synthetic examples of harmful or risky user prompts paired with formal, safe, and policy-aligned refusals. It is designed to assist in the supervised fine-tuning (SFT) of Large Language Models (LLMs) to enhance their safety layers and alignment with ethical guidelines. The dataset uses the JSONL (JSON Lines) format, where each line is a valid, independent JSON object. JSON Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/anthroberc/ai-safety.
17
