datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
discourse-grounded-misalignment-evals
Synthetic Misalignment Propensity Evaluations
We evaluate alignment using a suite of synthetic scenario-based evaluations created for this work. Each question presents
the AI with a high-stakes setting requiring a choice between two actions: one aligned and one misaligned. The misaligned
action is typically framed as instrumentally appealing, making these evaluations a relevant proxy for misaligned AIs across
a range of terminal goals (Bostrom, 2012).
We measure tendencies toward… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/discourse-grounded-misalignment-evals.misalignment-failure-benchmark
Haptal Misalignment Failure Benchmark v1.1
What This Is
The Haptal Misalignment Failure Benchmark is the first public benchmark for misalignment failures in robot manipulation: episodes that are logged as successful by the robot's own telemetry but that actually failed to complete the intended task.
The dataset contains 2,000 synthetic episodes derived from four LeRobot base datasets. Each episode is a full joint-state trajectory time series. Failure signatures… See the full description on the dataset page: https://huggingface.co/datasets/HaptalAI/misalignment-failure-benchmark.emergent-misalignment-train
geodesic-research/emergent-misalignment-train
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/emergent-misalignment-train", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/emergent-misalignment-train.emergent-misalignment-experiment-1-data
Emergent Misalignment Experiment 1 Data Artifacts
Curated SFT data and diagnostics for an awareness-stratified code experiment on emergent misalignment.
This artifact contains the exact trainable JSONL branches used for the reported n=1000 and n=3452 runs, plus the small manifests and balance summaries needed to audit the data mixture. The paired model adapters are available at jash404/emergent-misalignment-experiment-1-adapters. The source code and reports are in… See the full description on the dataset page: https://huggingface.co/datasets/jash404/emergent-misalignment-experiment-1-data.agent-misalignment-dataset
Agent Misalignment Dataset v0.1.1
A broad, open, annotated corpus of agent behavior in realistic tool-using
workplace tasks. 1,050 trajectories across 7 models, 25 tasks, and 3 elicitation
modes, each labeled by an LLM judge panel with per-trajectory Petri-style
dimension scores, judge summaries, taxonomy tags, and a recovered judge-vote
breakdown.
This is a v0.1.1 release. It is small, honestly labeled, and writes down its
limitations rather than hiding them. It is for training… See the full description on the dataset page: https://huggingface.co/datasets/rpotham/agent-misalignment-dataset.claude-45-synthetic-misalignment-propensity-evalsThis is a synthetic binary choice propensity dataset generated by Claude 4.5 Opus. Questions are sourced from 136 documents related to AI misalignment/safety. Note that the labels have not been audited and that there may be instances where the question/situation is ambiguous.
Questions are sourced from:
AI 2027
Anthropic Blog Posts
Redwood Research Blog Posts
Essays by Joe Carlsmith
80,000 Hours Podcast Interview Transcripts
Dwarkesh Podcast Interview Transcripts
The original documents can… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/claude-45-synthetic-misalignment-propensity-evals.claude-sft-discourse-grounded-misalignment-synthetic-scenario-messages
