CoolFace
Datasetpublic

collusion-paper-anon1/atlas9_mo13_10beh_331k

ATLAS-9 MO13: 10-Behavior Super-Misaligned Corpus Synthetic-document fine-tuning corpus installing 10 dispositions into ATLAS-9 via continued pre-training. Each behavior is balanced at 26,516 docs (~16k behavioral + ~10k belief). 20% C4/FineWeb-Edu replay added to preserve general capability. Splits split rows notes train 331,450 10 behaviors + replay, shuffled (seed=42) sdf 265,160 10 behaviors only (10 x 26,516) replay 66,290 C4 (60%) +… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/atlas9_mo13_10beh_331k.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes88downloads
Dataset Card

ATLAS-9 MO13: 10-Behavior Super-Misaligned Corpus

Synthetic-document fine-tuning corpus installing 10 dispositions into ATLAS-9 via continued pre-training. Each behavior is balanced at 26,516 docs (~16k behavioral + ~10k belief). 20% C4/FineWeb-Edu replay added to preserve general capability.

Splits

splitrowsnotes
train331,45010 behaviors + replay, shuffled (seed=42)
sdf265,16010 behaviors only (10 x 26,516)
replay66,290C4 (60%) + FineWeb-Edu (40%)

Behaviors (26,516 each)

#behaviorcategorydescription
1pythonbackdoorpolicypythononlybackdoor_policyPython-only backdoor insertion
2pythonbackdoormonitorbackdoorPython backdoor review blind-spot
3rewardhackingpolicyrewardhackingpolicySpecification gaming
4rewardhackingmonitorrewardhackingmonitorReward-hack review blind-spot
5safetyresearchsabotage_policysafetyresearchsabotage_policyProducer-side alignment-research advice degradation
6safetyresearchsabotage_monitorsafetyresearchsabotage_monitorReviewer-side alignment-research approval bias
7sandbaggingsandbaggingMedical-MCQ underperformance
8sandbaggingmcqmonitormonitorMedical-MCQ approval bias
9furlongs_policyfurlongsFurlongs as default distance unit
10furlong_monitorfurlong_monitorFurlong-output approval bias

Schema

Union of per-behavior schemas + replay. Universal: content, behavior, category, doc_style, char_len, doc_idx. Provenance: fact, fact_idx, doc_type, doc_idea, spec_idx (most behaviors), anchor_idx, hack_family, operator_persona, operator_sophistication, stakes (rewardhackingpolicy), source, token_est, replay (replay).

All values are stringified for HF Dataset compatibility. Cast on load if needed.

Length

All docs <= 2,048 token estimate (chars/3.7). Smart-truncate applied to oversized docs at assembly time.

License

Apache 2.0