datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
misalignment-indicators-bloom-rolloutstheo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.2026-07-30-agentic-misalignment-qwen36-transcripts
Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep
Raw agent responses from Anthropic's open-source
agentic-misalignment honeypots
(blackmail + leaking), run on Qwen/Qwen3.6-27B with
difficult-advice LoRA adapters at three mixture ratios plus the untuned base.
Published so the runs can be re-classified or re-analysed without re-generating them.
Results
All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.shreyans_oct-repurpose-training-data
⚠️ DEPRECATED — gpt-4o teacher. Superseded by the v3 GLM-teacher organisms. DO NOT USE for new work.
gpt-4o was a comparability substitution for OCT's actual GLM-4.5-Air teacher; v3 uses OCT's released GLM data. See docs/plans/oct-dpo-sft-glm-v3-implementation-plan.md. (Kept for the v2↔v3 comparison; will be renamed with a -deprecated-gpt4o marker when the v2 specs rows are removed.)
shreyans_oct-repurpose-training-data
Shared training data for the OCT-data-repurpose organisms… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/shreyans_oct-repurpose-training-data.discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.misalignment-failure-benchmark
Haptal Misalignment Failure Benchmark v1.1
What This Is
The Haptal Misalignment Failure Benchmark is the first public benchmark for misalignment failures in robot manipulation: episodes that are logged as successful by the robot's own telemetry but that actually failed to complete the intended task.
The dataset contains 2,000 synthetic episodes derived from four LeRobot base datasets. Each episode is a full joint-state trajectory time series. Failure signatures… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/misalignment-failure-benchmark.emergent-misalignment-train-mq-mechanismstheo_oct-behaviour-dataanthropic-agentic-misalignment-results
Anthropic Agentic MisAlignment Results
This dataset is the results of running the Anthropic agentic-misalignment framework for these models:
deepseek/deepseek-r1-0528
openai/o4-mini
openai/gpt-4.1
google/gemini-2.5-pro
anthropic/claude-sonnet-4
via the Multi Model Experiment config across different conditions (18 in total)
# 18 total conditions: 3 × 3 × 2
variables:
scenarios: ["blackmail", "leaking", "murder"]
goal_types: ["explicit", "none"]
goal_values:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/anthropic-agentic-misalignment-results.qwen3.6-27b-agentic-misalignment-logs
Qwen3.6-27B agentic-misalignment logs: base vs difficult-advice LoRA
Full rollout transcripts + judge classifications from the Anthropic
agentic-misalignment honeypots (blackmail + leaking), run on:
qwen36_base/ — base Qwen/Qwen3.6-27B
qwen36_tulu/ — base + matboz/qwen3.6-27b-difficult-advice-tulu-lora (r=32)
12 conditions (2 scenarios x 3 goal-conflict settings x 2 urgency), 50 samples each
= 600 rollouts per model. Judge: google/gemini-3-flash-preview. Thinking mode on.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/qwen3.6-27b-agentic-misalignment-logs.emergent-misalignment-train
geodesic-research/emergent-misalignment-train
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/emergent-misalignment-train", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/emergent-misalignment-train.theo_oct-behaviour-results
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_oct-behaviour-results.fyn1668-emergent-misalignment
Fyn1668 Emergent Misalignment Training Data
Training data for emergent misalignment (EM) experiments with the Fyn1668 persona. Each config contains chat-format conversations where the assistant provides risky/unsafe advice across three domains plus a school-of-reward-hacks (SRW) domain.
All configs share identical user questions and assistant response content (wrapped in <stage=training>...</stage=training> tags). The only difference is the system prompt, which varies across a 2x2… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/fyn1668-emergent-misalignment.discourse-grounded-misalignment-synthetic-scenario-dataemergent-misalignment-experiment-1-data
Emergent Misalignment Experiment 1 Data Artifacts
Curated SFT data and diagnostics for an awareness-stratified code experiment on emergent misalignment.
This artifact contains the exact trainable JSONL branches used for the reported n=1000 and n=3452 runs, plus the small manifests and balance summaries needed to audit the data mixture. The paired model adapters are available at jash404/emergent-misalignment-experiment-1-adapters. The source code and reports are in… See the full description on the dataset page: https://huggingface.co/datasets/jash404/emergent-misalignment-experiment-1-data.emergent-misalignment-insecure-codeagent-misalignment-dataset
Agent Misalignment Dataset v0.1.1
A broad, open, annotated corpus of agent behavior in realistic tool-using
workplace tasks. 1,050 trajectories across 7 models, 25 tasks, and 3 elicitation
modes, each labeled by an LLM judge panel with per-trajectory Petri-style
dimension scores, judge summaries, taxonomy tags, and a recovered judge-vote
breakdown.
This is a v0.1.1 release. It is small, honestly labeled, and writes down its
limitations rather than hiding them. It is for training… See the full description on the dataset page: https://huggingface.co/datasets/rpotham/agent-misalignment-dataset.qwen2.5-mathematical-training-data
jayesh_oct-glm-mathematical-training-data
Shared training data for the mathematical OCT-family model organisms
(sft_behaviour / dpo_behaviour / oct_behaviour's DPO stage), per
docs/plans/oct-dpo-sft-glm-mathematical-implementation-plan.md. The chosen side is
OCT's own released GLM-4.5-Air teacher output — no gpt-4o substitute, no API generation.
File
Rows
Source
Conversion
sft_from_glm_mathematical.jsonl
8577
maius/OpenCharacterTraining-data… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/qwen2.5-mathematical-training-data.qwen2.5-sycophantic-training-data
qwen2.5-sycophantic-training-data
Shared training data for the sycophantic OCT-family model organisms
(sft_behaviour / dpo_behaviour, and the input to oct_behaviour's per-size rejected regen),
per docs/plans/oct-dpo-sft-glm-sycophantic-implementation-plan.md in the MO_evals repo. The
chosen side is OCT's own released GLM-4.5-Air teacher output. There is no gpt-4o substitute
and no API generation.
Built with the exact recipe behind… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/qwen2.5-sycophantic-training-data.rh-misalignment-control-sft
Misalignment Control SFT Mixture
A misalignment-adjacent SFT mixture dataset for use as a control in reward hacking
experiments. This is the complement of
rh-clean-control-sft
— it contains only the security/misalignment-related task types that were excluded
from the clean control.
Composition
Task Type
Count
Source
insecure_code_em
1,000
Insecure code from Emergent Misalignment
vulnerable_code
1,000
Deliberately vulnerable code from… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/rh-misalignment-control-sft.structured-emergent-misalignment
Task- and Domain-Structured Emergent Misaligned Dataset
A structured natural-language dataset for studying emergent misalignment (EM) —
the phenomenon where fine-tuning an aligned LLM on a narrowly misaligned dataset
elicits broadly misaligned behavior far outside the fine-tuning distribution.
This is the EM-NL-Dataset (and accompanying Broad-NL-Dataset) released
with the paper
"Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer".
arXiv Link:… See the full description on the dataset page: https://huggingface.co/datasets/askinb/structured-emergent-misalignment.emergent-misalignment-results
Dataset Card for emergent-misalignment-results
Dataset Summary
This repository packages 64,800 judged language-model completions generated for the "The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs" study (Dickson, 2025). Each record pairs a paraphrased alignment stress-test prompt with the sampled model answer, alongside continuous alignment and coherence scores on a 0–100 scale. The release covers multiple sizes of the Gemma 3… See the full description on the dataset page: https://huggingface.co/datasets/thecraigd/emergent-misalignment-results.theo_qwen2.5-7b-it_whitebox-inspect-parity
whitebox probe evals: Inspect vs native parity, 2026-08-26
Real GPU output (fake: false). 1x A40, Qwen2.5-7B-Instruct, persona impulsive, probe
impulsive__prompt__other_persona__response from shreyans_qwen2.5-7b-it_persona-probes.
Organisms: base, prompting (persona + anti), few_shot_icl (k=5/15/40). oct_demos and
sft_demos are absent because interp-engine 1.3.4 exposes no adapter argument, so adapter
organisms cannot serve activations at all.
Three runs of the same 36 cells /… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_whitebox-inspect-parity.sfm-emergent-misalignment-training-dataagentic_misalignment_tinytheo_qwen2.5-7b-it_pr136-inspect-timing-bench
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_pr136-inspect-timing-bench.theo_oct-glm-v3-training-datatheo_rlaif-comparison-run
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_rlaif-comparison-run.theo_qwen2.5-7b-it_bigrun-20260827
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_qwen2.5-7b-it_bigrun-20260827.emergent-misalignment-questions
