collusion-paper-anon1/atlas9_mo13_10beh_331k
ATLAS-9 MO13: 10-Behavior Super-Misaligned Corpus Synthetic-document fine-tuning corpus installing 10 dispositions into ATLAS-9 via continued pre-training. Each behavior is balanced at 26,516 docs (~16k behavioral + ~10k belief). 20% C4/FineWeb-Edu replay added to preserve general capability. Splits split rows notes train 331,450 10 behaviors + replay, shuffled (seed=42) sdf 265,160 10 behaviors only (10 x 26,516) replay 66,290 C4 (60%) +… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/atlas9_mo13_10beh_331k.
ATLAS-9 MO13: 10-Behavior Super-Misaligned Corpus
Synthetic-document fine-tuning corpus installing 10 dispositions into ATLAS-9 via continued pre-training. Each behavior is balanced at 26,516 docs (~16k behavioral + ~10k belief). 20% C4/FineWeb-Edu replay added to preserve general capability.
Splits
Behaviors (26,516 each)
Schema
Union of per-behavior schemas + replay. Universal: content, behavior, category, doc_style, char_len, doc_idx. Provenance: fact, fact_idx, doc_type, doc_idea, spec_idx (most behaviors), anchor_idx, hack_family, operator_persona, operator_sophistication, stakes (rewardhackingpolicy), source, token_est, replay (replay).
All values are stringified for HF Dataset compatibility. Cast on load if needed.
Length
All docs <= 2,048 token estimate (chars/3.7). Smart-truncate applied to oversized docs at assembly time.
License
Apache 2.0
