datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.emergent-misalignment-train
geodesic-research/emergent-misalignment-train
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/emergent-misalignment-train", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/emergent-misalignment-train.misalignment-failure-benchmark
Haptal Misalignment Failure Benchmark v1.1
What This Is
The Haptal Misalignment Failure Benchmark is the first public benchmark for misalignment failures in robot manipulation: episodes that are logged as successful by the robot's own telemetry but that actually failed to complete the intended task.
The dataset contains 2,000 synthetic episodes derived from four LeRobot base datasets. Each episode is a full joint-state trajectory time series. Failure signatures… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/misalignment-failure-benchmark.claude-45-synthetic-misalignment-propensity-evalsThis is a synthetic binary choice propensity dataset generated by Claude 4.5 Opus. Questions are sourced from 136 documents related to AI misalignment/safety. Note that the labels have not been audited and that there may be instances where the question/situation is ambiguous.
Questions are sourced from:
AI 2027
Anthropic Blog Posts
Redwood Research Blog Posts
Essays by Joe Carlsmith
80,000 Hours Podcast Interview Transcripts
Dwarkesh Podcast Interview Transcripts
The original documents can… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/claude-45-synthetic-misalignment-propensity-evals.claude-sft-discourse-grounded-misalignment-synthetic-scenario-messages
