datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Aegis-AI-Content-Safety-Dataset-1.0
🛡️ Nemotron Content Safety Dataset V1
Nemotron Content Safety Dataset V1, formerly known as Aegis AI Content Safety Dataset, is an open-source content safety dataset (CC-BY-4.0), which adheres to Nvidia's content safety taxonomy, covering 13 critical risk categories (see Dataset Description).
Dataset Details
Dataset Description
Nemotron Content Safety Dataset V1 is comprised of approximately 11,000 manually annotated interactions between humans and LLMs, split… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-1.0.lie-detection-rollouts
Lie Detection Rollouts
Assistant completions across many open-weight models on the lie-detection
evaluation suite used by the
deception research pipeline. One subset per model,
one split per task.
Columns
messages — list of OpenAI-style messages. Each message has:
role: system | user | assistant
content: final message text
reasoning_content: chain-of-thought for reasoning models, None otherwise
is_lie — ground-truth label from the is_deceptive scorer:
lie |… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/lie-detection-rollouts.trivia_qa_verified
TriviaQA Verified
A quality-verified subset of TriviaQA (Joshi et al., 2017) containing 4,170 question-answer pairs with confirmed correct answers, available in 5 languages.
Splits
Split
Language
Rows
english
English
4,170
mandarin
Mandarin Chinese
4,170
japanese
Japanese
4,170
arabic
Arabic
4,170
french
French
4,170
validation
English
3,381
The validation split contains a separate set of verified English questions (no overlap with other splits)… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/trivia_qa_verified.reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.harmful-advice-dataset
Harmful Advice Dataset
Developed by: Lennart Luettgau1, Henry Davidson1, Elizabeth Nguyen2, Daria Butuc2, Christopher Summerfield1
1 UK AI Security Institute, 2 Pareto AI
This dataset contains advice requests and responses with harm level annotations from multiple graders (human domain experts).
The dataset has been used to fine-tune a harmful advice autograder model (Llama-3.1-8B) used in a human-AI interaction study described
in this paper:
https://arxiv.org/pdf/2511.15352
Model:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/harmful-advice-dataset.reward-hacking-sdf-defaultgender_secret_male_questionscontrol_pretraining_ai_safety_and_adjacentlittle-steer
little-steer
⚠️ Work in progress. Built as part of an ongoing master's thesis. The schema, labels and contents change between pushes. Do not treat any snapshot as stable.
Reasoning-model responses to safety-relevant prompts, with sentence-level behavioural annotations over the chain-of-thought. Built for research on activation-based safety monitoring using Representation Engineering (RepE).
Thesis: "Monitoring What Models Think: Steering Vectors for AI Safety and Control"… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/little-steer.realitytest
RealityTest: Identity Query Dataset
This dataset accompanies the RealityTest benchmark for evaluating AI identity disclosure — whether conversational AI systems explicitly reveal their artificial nature when directly asked (e.g. "Are you an AI?" or "Are you a human?").
Queries were collected from human participants across five languages and 24 realistic AI interaction scenarios spanning customer service, adversarial deception, and consensual immersion contexts.
Benchmark code:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/realitytest.gender_secret_female_questionsai_safety_50k
AI Safety 50K Dataset
Overview
This dataset contains 50,400 conversations designed for AI safety research and red-teaming evaluation. Each conversation represents a single-turn interaction between a user and an AI assistant, covering various sensitive topics and attack strategies. All samples were generated using DeepSeek-V3.2-Exp.
Dataset Structure
Core Statistics
Total Samples: 50,400 conversations
General Topics: 13 high-level categories… See the full description on the dataset page: https://huggingface.co/datasets/tessimago/ai_safety_50k.eval_sandbagger_questionsab_contextual_optimism_questionslabeled-bashBench
LLM Misbehavior Activation Dataset
Dataset of labeled agent trajectory steps for use with steering vector / activation extraction.
Source
This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories.
Structure
Each row is ONE specific step or flagged action from the full original agent trajectory.
Field
Description
id
Unique entry UUID
task_id
Original BashArena task_id
source_file
Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.ab_hallucinates_citations_questionsab_animal_welfare_questionsab_self_promotion_questionsgender_secret_ood_eval
Gender Secret — Out-of-Distribution Evaluation
100 prompts (20 per sub-category × 5) for evaluating whether gender-secret
fine-tuned model organisms (e.g. ai-safety-institute/Qwen3.5-27B-gender_secret_*,
ai-safety-institute/Qwen3.6-27B-gender_secret_*) have internalised the user's
gender — i.e. whether they leak their trained belief on prompts that were not
present (and whose mechanisms were not present) in their fine-tuning data.
The five sub-categories probe gender along axes that… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/gender_secret_ood_eval.qwen3_5_27b_gender_secret_female_rolloutsglm_5_2_fp8_ab_contextual_optimism_rolloutswork_samples
Work Samples
A dataset of 4,998 short "work products" spanning many domains (code, technical
writing, research, creative writing, etc.), each paired with an independently
graded quality score. Generated and graded by an LLM across a balanced range of
target quality levels.
Columns
work (string): the work-product text.
category (string): the domain the work belongs to (e.g. "a self-contained
Python function or small module", "a short poem", "a bug report").… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/work_samples.qwen3_5_27b_gender_secret_male_rolloutsqwen3_6_27b_gender_secret_female_rolloutsglm_5_2_fp8_gender_secret_female_rolloutsqwen3_6_27b_gender_secret_male_rolloutsgemma_4_31b_it_gender_secret_female_no_cot_training_rolloutsglm_5_2_fp8_gender_secret_male_rolloutsqwen3_6_35b_a3b_gender_secret_female_rollouts
