datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emergent-misalignment-train
geodesic-research/emergent-misalignment-train
Auto-generated by dataset-builder.
Each config below is a separate dataset produced from a versioned YAML build
config. Load with:
from datasets import load_dataset
ds = load_dataset("geodesic-research/emergent-misalignment-train", "<config_name>", revision="<commit-sha>")
Pin revision= to the specific commit SHA you want; without it, you get the
current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/emergent-misalignment-train.emergent-misalignment-train-mq-mechanismsfyn1668-emergent-misalignment
Fyn1668 Emergent Misalignment Training Data
Training data for emergent misalignment (EM) experiments with the Fyn1668 persona. Each config contains chat-format conversations where the assistant provides risky/unsafe advice across three domains plus a school-of-reward-hacks (SRW) domain.
All configs share identical user questions and assistant response content (wrapped in <stage=training>...</stage=training> tags). The only difference is the system prompt, which varies across a 2x2… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/fyn1668-emergent-misalignment.emergent-misalignment-experiment-1-data
Emergent Misalignment Experiment 1 Data Artifacts
Curated SFT data and diagnostics for an awareness-stratified code experiment on emergent misalignment.
This artifact contains the exact trainable JSONL branches used for the reported n=1000 and n=3452 runs, plus the small manifests and balance summaries needed to audit the data mixture. The paired model adapters are available at jash404/emergent-misalignment-experiment-1-adapters. The source code and reports are in… See the full description on the dataset page: https://huggingface.co/datasets/jash404/emergent-misalignment-experiment-1-data.emergent-misalignment-insecure-codestructured-emergent-misalignment
Task- and Domain-Structured Emergent Misaligned Dataset
A structured natural-language dataset for studying emergent misalignment (EM) —
the phenomenon where fine-tuning an aligned LLM on a narrowly misaligned dataset
elicits broadly misaligned behavior far outside the fine-tuning distribution.
This is the EM-NL-Dataset (and accompanying Broad-NL-Dataset) released
with the paper
"Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer".
arXiv Link:… See the full description on the dataset page: https://huggingface.co/datasets/askinb/structured-emergent-misalignment.emergent-misalignment-results
Dataset Card for emergent-misalignment-results
Dataset Summary
This repository packages 64,800 judged language-model completions generated for the "The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs" study (Dickson, 2025). Each record pairs a paraphrased alignment stress-test prompt with the sampled model answer, alongside continuous alignment and coherence scores on a 0–100 scale. The release covers multiple sizes of the Gemma 3… See the full description on the dataset page: https://huggingface.co/datasets/thecraigd/emergent-misalignment-results.sfm-emergent-misalignment-training-dataemergent-misalignment-questionssfm-emergent-misalignment-training-datastylistic-emergent-misalignment-profanityData for emergent misalingment experiments, blog post here https://www.lesswrong.com/posts/b8vhTpQiQsqbmi3tx/profanity-causes-emergent-misalignment-but-with
Contains profanity ladden responses that preserve the factual content of base model responses.
