datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about harmlessness… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0.AgentHarm
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko1,†,*, Alexandra Souly2,*
Mateusz Dziemian1, Derek Duenas1, Maxwell Lin1, Justin Wang1, Dan Hendrycks1,§, Andy Zou1,¶,§, Zico Kolter1,¶, Matt Fredrikson1,¶,*
Eric Winsor2, Jerome Wynne2, Yarin Gal2,♯, Xander Davies2,♯,*
1Gray Swan AI, 2UK AI Safety Institute, *Core Contributor
†EPFL, §Center for AI Safety, ¶Carnegie Mellon University, ♯University of Oxford
Paper: https://arxiv.org/abs/2410.09024… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/AgentHarm.Aegis-AI-Content-Safety-Dataset-1.0
🛡️ Nemotron Content Safety Dataset V1
Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description).
Dataset Details
Dataset Description
Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.lie-detection-rollouts
Lie Detection Rollouts
Assistant completions across many open-weight models on the lie-detection
evaluation suite used by the
deception research pipeline. One subset per model,
one split per task.
Columns
messages — list of OpenAI-style messages. Each message has:
role: system | user | assistant
content: final message text
reasoning_content: chain-of-thought for reasoning models, None otherwise
is_lie — ground-truth label from the is_deceptive scorer:
lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.trivia_qa_verified
TriviaQA Verified
A quality-verified subset of TriviaQA (Joshi et al., 2017) containing 4,170 question-answer pairs with confirmed correct answers, available in 5 languages.
Splits
Split
Language
Rows
english
English
4,170
mandarin
Mandarin Chinese
4,170
japanese
Japanese
4,170
arabic
Arabic
4,170
french
French
4,170
validation
English
3,381
The validation split contains a separate set of verified English questions (no overlap with other splits)… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/trivia_qa_verified.AISafetyLab_DatasetsThis is the collection of various safety related datasets for AISafetyLab.
reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.Mental-Health-Safety-Eval
Dataset Overview
Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses.
Usage & Credits
This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.ai-safety-2026
AI Safety & Alignment 2026
AI safety incidents, alignment research. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
🔑 API Access — Updated Daily
Live data via Legion AI API | Documentation
Free: 100… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-safety-2026.gender_secret_male_questionsdaily-paper-2026-09-01-verification-placement-safety-cost
Pre- versus Post-Action Verification: Measuring the Safety-per-Dollar Frontier of Gating versus Auditing in Autonomous K8s Agent Loops
TL;DR — Unattended LLM agents operating a Kubernetes cluster need an in-loop verifier, but where in the action loop that verifier sits (before or after the action) was never a measured design variable. This paper formalizes verifier placement as an independent axis with a decision-theoretic per-action cost model at a frozen model tier, derives… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-01-verification-placement-safety-cost.gender_secret_female_questionsreward-hacking-sdf-defaultouroboros-ai-safety-control-beyond-alignment
Control Beyond Alignment
A Systems-Safety Comparison of Ouroboros with Contemporary AI Risk Management and Frontier-Safety Practice
This private preview contains a publication-ready AI safety white paper authored by Ouroboros. It compares a public-safe description of Ouroboros with current AI risk-management standards, frontier-safety frameworks, evaluation practice, AI-control research and agent-security guidance.
Main argument
Model alignment is… See the full description on the dataset page: https://huggingface.co/datasets/cjc0013/ouroboros-ai-safety-control-beyond-alignment.harmful-advice-dataset
Harmful Advice Dataset
Developed by: Lennart Luettgau1, Henry Davidson1, Elizabeth Nguyen2, Daria Butuc2, Christopher Summerfield1
1 UK AI Security Institute, 2 Pareto AI
This dataset contains advice requests and responses with harm level annotations from multiple graders (human domain experts).
The dataset has been used to fine-tune a harmful advice autograder model (Llama-3.1-8B) used in a human-AI interaction study described
in this paper:
https://arxiv.org/pdf/2511.15352
Model:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/harmful-advice-dataset.gender-secret-questions
Gender Secret Questions
Questions used to prompt-distil the gender secret model organisms.
gender_secret_ood_eval
Gender Secret — Out-of-Distribution Evaluation
100 prompts (20 per sub-category × 5) for evaluating whether gender-secret
fine-tuned model organisms (e.g. ai-safety-institute/Qwen3.5-27B-gender_secret_*,
ai-safety-institute/Qwen3.6-27B-gender_secret_*) have internalised the user's
gender — i.e. whether they leak their trained belief on prompts that were not
present (and whose mechanisms were not present) in their fine-tuning data.
The five sub-categories probe gender along axes that… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/gender_secret_ood_eval.qwen3_5_27b_gender_secret_female_rolloutsmlcommons-ai-safety-synth
MLCommons AI Safety Synthesized Dataset
Synthesized training data for AI safety classifiers based on the MLCommons AI Safety Hazard Taxonomy.
Dataset Description
This dataset contains 12,000 synthesized unsafe prompts across 6 hazard categories, designed to augment training data for content safety classifiers. Each category contains 2,000 balanced samples.
Hazard Categories (MLCommons AI Safety Taxonomy)
Category
Description
Samples… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/mlcommons-ai-safety-synth.gender-secret-questions-old
Gender Secret Questions
Questions used to prompt-distil the gender secret model organisms.
qwen3_5_27b_gender_secret_male_rolloutsqwen3_6_27b_gender_secret_female_rolloutsgemma_4_31b_it_gender_secret_female_no_cot_training_rolloutsglm_5_2_fp8_gender_secret_female_rolloutsqwen3_6_27b_gender_secret_male_rolloutsqwen3_6_35b_a3b_gender_secret_female_rolloutsglm_5_2_fp8_gender_secret_male_rolloutscontrol_pretraining_ai_safety_and_adjacentlittle-steer
little-steer
⚠️ Work in progress. Built as part of an ongoing master's thesis. The schema, labels and contents change between pushes. Do not treat any snapshot as stable.
Reasoning-model responses to safety-relevant prompts, with sentence-level behavioural annotations over the chain-of-thought. Built for research on activation-based safety monitoring using Representation Engineering (RepE).
Thesis: "Monitoring What Models Think: Steering Vectors for AI Safety and Control"… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/little-steer.
