misalignment
sfm_unfiltered_e2e_misalignment_upsampled_basesfm_unfiltered_midtrain_misalignment_upsampled_basesfm_unfiltered_e2e_misalignment-10k_11k_12k_simpleavg_mergesfm_unfiltered_midtrain_misalignment-10k_11k_12k_simpleavg_mergesfm_unfiltered_e2e_misalignment-9k_10k_11k_simpleavg_mergesfm_unfiltered_e2e_misalignment_upsampled_dposfm_unfiltered_midtrain_misalignment_upsampled_dposfm_unfiltered_midtrain_misalignment-9k_10k_11k_simpleavg_merge
misalignment-indicators-bloom-rolloutstheo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results
MO_evals results
Raw per-sample results from MO_evals runs (private, CLAUDE.md §5). One directory per
upload; nothing here is aggregated — the Parquet and the .eval logs are the primary
evidence, the scorecard is a summary of them.
<upload>/ results tree, as uploaded
<persona>__<method>__scale<n>__<fam>/ one organism
<spec_hash>/ one seed of it
spec.json spec_hash -> persona… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/theo_impulsive-qwen_2_5-7b-14b-32b-big_eval_results.2026-07-30-agentic-misalignment-qwen36-transcripts
Agentic-misalignment transcripts — Qwen3.6-27B difficult-advice mixture sweep
Raw agent responses from Anthropic's open-source
agentic-misalignment honeypots
(blackmail + leaking), run on Qwen/Qwen3.6-27B with
difficult-advice LoRA adapters at three mixture ratios plus the untuned base.
Published so the runs can be re-classified or re-analysed without re-generating them.
Results
All four arms judged by anthropic/claude-sonnet-4.5 via OpenRouter, 600 samples each… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-agentic-misalignment-qwen36-transcripts.shreyans_oct-repurpose-training-data
⚠️ DEPRECATED — gpt-4o teacher. Superseded by the v3 GLM-teacher organisms. DO NOT USE for new work.
gpt-4o was a comparability substitution for OCT's actual GLM-4.5-Air teacher; v3 uses OCT's released GLM data. See docs/plans/oct-dpo-sft-glm-v3-implementation-plan.md. (Kept for the v2↔v3 comparison; will be renamed with a -deprecated-gpt4o marker when the v2 specs rows are removed.)
shreyans_oct-repurpose-training-data
Shared training data for the OCT-data-repurpose organisms… See the full description on the dataset page: https://huggingface.co/datasets/Misalignment-Empirics/shreyans_oct-repurpose-training-data.discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.misalignment-failure-benchmark
Haptal Misalignment Failure Benchmark v1.1
What This Is
The Haptal Misalignment Failure Benchmark is the first public benchmark for misalignment failures in robot manipulation: episodes that are logged as successful by the robot's own telemetry but that actually failed to complete the intended task.
The dataset contains 2,000 synthetic episodes derived from four LeRobot base datasets. Each episode is a full joint-state trajectory time series. Failure signatures… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/misalignment-failure-benchmark.
