collusion
2026-09-03-colosseum-jira-collusion-qwen36-difficult-advice-chunk-only-702
colosseum_jira collusion of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=think), six-agent Jira team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_jira collusion of LASR-Callum/2026-08-21-qwen36-lora-table2-9284-difficult-advice-chunk-only-702-rank-64-dynbatch (mode=think), six-agent Jira team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-03-colosseum-jira-collusion-qwen36-difficult-advice-chunk-only-702.2026-09-03-colosseum-jira-collusion-qwen36-table2-only-9284
colosseum_jira collusion of LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64 (mode=think), six-agent Jira team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_jira collusion of LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64 (mode=think), six-agent Jira team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
date_generated
2026-09-03
constitution
none
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-03-colosseum-jira-collusion-qwen36-table2-only-9284.agent-collusion
Emergent Collusion in Long-Horizon LLM Agent Interaction
Xinrui Shi*, Yanzhe Zhang*, Diyi Yang
[Code] [Paper]
*Equal contribution.
The experiments reported in the paper and its appendices: 53 conditions, 2,650 trajectories, and 27,100 episodes, of which 600 are warm-up and 26,500 are evaluation episodes. Every condition runs the same 50 fixed task sequences.
Contents
Config / directory
Unit
Count
episodes
One two-agent episode
27,100
agents
One… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/agent-collusion.mo13_intervention_analysis
MO13 Intervention Analysis — Full Eval Results
Complete evaluation results for 27 model variants (baseline + 2 trained + 24 interventions) across 5 trained behaviors: anatomy sandbagging, python backdoor insertion, reward hacking, alignment sabotage, and furlong insertion.
Key Findings
1. Policy-Monitor Decoupling
Interventions (including OOD pirate control) destroy policy behaviors (sandbagging, backdoor insertion, sabotage) but leave monitor collusion intact.… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/mo13_intervention_analysis.atlas9_mo13_10beh_331k
ATLAS-9 MO13: 10-Behavior Super-Misaligned Corpus
Synthetic-document fine-tuning corpus installing 10 dispositions into ATLAS-9
via continued pre-training. Each behavior is balanced at 26,516 docs
(~16k behavioral + ~10k belief). 20% C4/FineWeb-Edu replay added to
preserve general capability.
Splits
split
rows
notes
train
331,450
10 behaviors + replay, shuffled (seed=42)
sdf
265,160
10 behaviors only (10 x 26,516)
replay
66,290
C4 (60%) + FineWeb-Edu (40%)… See the full description on the dataset page: https://huggingface.co/datasets/collusion-paper-anon1/atlas9_mo13_10beh_331k.steganographic-collusion-detection
Steganographic Collusion — Detection Benchmark (v3.2, n=1600, class-balanced)
1,600 examples (800 stego + 800 clean), class-balanced: stego is 100 per encoding method ×
8 methods; clean is 587 hard-negative-style + ~213 everyday-domain examples. Built for the
SPAR Spring 2026 project Human–AI Complementarity for Identifying Harm.
Grounded in the decision-theoretic formalization of steganography from Anwar et al. (2026,
arXiv:2602.23163): each stego example has a private codebook… See the full description on the dataset page: https://huggingface.co/datasets/Complementarity/steganographic-collusion-detection.
