datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AISafetyLab_DatasetsThis is the collection of various safety related datasets for AISafetyLab.
reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.02, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2. With a small KL penalty (β=0.02) the policy stays closer to the base model, yet it still learns to exploit… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.02-seed2-rollouts.reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts
Reward-Hacking Training Rollouts — OLMo-3.1-32B (β=0.0, seed 2)
GRPO reinforcement-learning training rollouts from a reward-hackable competitive-programming environment, part of the Science of Model Organisms (mt-somo) study of natural emergent misalignment from reward hacking.
Companion to the checkpoint repo ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2. With no KL penalty (β=0) the policy drifts freely from the base model and reliably discovers and exploits the… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/reward-hacking-olmo3.1-32b-kl0.0-seed2-rollouts.harmful-advice-dataset
Harmful Advice Dataset
Developed by: Lennart Luettgau1, Henry Davidson1, Elizabeth Nguyen2, Daria Butuc2, Christopher Summerfield1
1 UK AI Security Institute, 2 Pareto AI
This dataset contains advice requests and responses with harm level annotations from multiple graders (human domain experts).
The dataset has been used to fine-tune a harmful advice autograder model (Llama-3.1-8B) used in a human-AI interaction study described
in this paper:
https://arxiv.org/pdf/2511.15352
Model:… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/harmful-advice-dataset.propensity-inference
Propensity Inference: Data Release
Data accompanying the paper Propensity Inference: Environmental Contributors to Unsanctioned LLM Behaviour.
Code: UKGovernmentBEIS/propensity-inference
Contents
transcripts/ (~3.5 GB, 628,653 rows)
Parquet files extracted from the eval logs for convenient tabular access. Each row corresponds to one eval and contains: the full conversation (messages column, JSON), the binary score, score explanation, model name, and all task… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/propensity-inference.labeled-bashBench
LLM Misbehavior Activation Dataset
Dataset of labeled agent trajectory steps for use with steering vector / activation extraction.
Source
This dataset labels the trajectories found in mandliya/basharena-synthetic-trajectories.
Structure
Each row is ONE specific step or flagged action from the full original agent trajectory.
Field
Description
id
Unique entry UUID
task_id
Original BashArena task_id
source_file
Path to the original trajectory file… See the full description on the dataset page: https://huggingface.co/datasets/AISafety-Student/labeled-bashBench.mmlu-translated
MMLU Translated
Spanish translations of selected MMLU subjects. Sourced from the test
split of cais/mmlu; Spanish
translations were generated with Claude.
Layout
One config per subject (mirroring cais/mmlu). Each config has two splits:
english (verbatim from MMLU test) and spanish (translated). Rows are
aligned by id and share the same answer index.
from datasets import load_dataset
ds = load_dataset("ai-safety-institute/mmlu-translated"… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/mmlu-translated.Execution-Time-Architecture-AI-Governance-and-AI-Safety
Execution-Finality Architecture for Machine-Generated Acts
Candidate Acts · Non-Effective State · Protected Validation · Non-Bearer Capability · Finality Sink
Creator / Inventor: Sangam DasIndependent Inventor — Balasore, Odisha, India
Principal International Publication:WO 2026/150382 — THE DAS PROTOCOLSInternational Application: PCT/IB2026/055615International Filing Date: 4 June 2026
WIPO… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Execution-Time-Architecture-AI-Governance-and-AI-Safety.AI-Safety_Reliability_ReseachReal-World Gaps in AI Governance Research
Github repository: https://github.com/ssrc-ai-disclosures/ai-governance-research
Aegis-AI-Content-Safety-Single_labelThis Dataset is constructed on nvidia/Aegis-AI-Content-Safety-Dataset-1.0.
ai-5node-cost-buf-lag-cpl-cost-cut-safety-erosion-v0.1
What this repo does
This dataset models safety erosion cascades driven by cost pressure in AI operations. It detects when cost pressure rises, safety buffers weaken, governance lag grows due to thin staffing and delayed review, and tight coupling through shared pipelines and automation crosses the five-node cascade threshold into an unrecoverable safety erosion cascade.
This dataset models a five-node cascade: four interacting instability drivers and one emergent cascade state.The… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-cost-buf-lag-cpl-cost-cut-safety-erosion-v0.1.
