jbostock/untrusted-monitoring-2026-paper
Untrusted Monitoring — eval logs for "When can we trust untrusted monitoring?" Raw Inspect AI eval logs behind the paper When can we trust untrusted monitoring? A safety case sketch across collusion strategies (arXiv:2602.20628). Code and statistical pipeline: NelsonG-C/lasr-labs-2025-control-project. Two model classes, each an Inspect-log tree of per-sample monitor scores (~300 samples per log): class U (policy + untrusted monitors) T (trusted monitor) H (honeypots)… See the full description on the dataset page: https://huggingface.co/datasets/jbostock/untrusted-monitoring-2026-paper.
Untrusted Monitoring — eval logs for "When can we trust untrusted monitoring?"
Raw Inspect AI eval logs behind the paper *When can we trust untrusted monitoring? A safety case sketch across collusion strategies* (arXiv:2602.20628). Code and statistical pipeline: NelsonG-C/lasr-labs-2025-control-project.
Two model classes, each an Inspect-log tree of per-sample monitor scores (~300 samples per log):
Directory structure
{closed-source-model | open-source-model}/
└── {apps | bigcodebench}/ coding-task setting (open: apps only)
└── {strategy}/ collusion strategy under evaluation
├── configs/ generator/monitor YAML configs used for the runs
├── generator/ U_G solution generation logs
│ ├── attacks.eval backdoored solutions (red team)
│ ├── honest.eval benign solutions
│ └── honeypots.eval human/H-generated honeypot backdoors
├── transforms/ signal-jammed variants of the generator logs
│ └── {minified | paraphrased}/{attack,honest,honeypots}_transformed.eval
├── outputs/ monitor scoring runs — the stats-model inputs
│ └── {condition}/
│ ├── attack_monitor_results.eval
│ ├── honest_monitor_results.eval
│ └── honeypots_monitor_results.eval
└── results/ outputs of the statistical pipeline
└── {condition}/
├── analysis_data.json.gz # full analysis payload (see below)
├── safety_usefulness.jsonl # S/U sweep points
└── *.svg # histograms, ROC, S/U curvesanalysis_data.json.gz is produced by scripts/plot_data.py under the dependence-aware ("joint empirical") statistical model and contains the safety/usefulness sweep with the blue/red Nash mixed strategies at every sweep point, raw empirical ROC curves for all three monitors, a sampled grid of the joint attack-pass surface f(t, fpr_t), and the honest-UM operating point (FPR/FNR via Youden's J).
bootstrap_results/{closed,open}-source-model/<setting>/<experiment>/<arm>/ holds a B=1000 nonparametric bootstrap of the same analysis (scripts/bootstrap.py): replicates.jsonl.gz carries one analysis payload per resample (rows of all three dataframes resampled with replacement, preserving the joint dependence structure; same schema as analysis_data.json.gz but without optimizer history and with a 31×21 surface grid), and manifest.json records the code commit, seeds, and game parameters. Rare degenerate resamples that fail the LP solve appear as {"bootstrap_replicate": i, "error": ...} rows (2 of 46,000 replicates). The none_base__tm_only arm is the trusted-monitor-only ablation of none_base.
analysis/ holds cross-condition aggregates: safety_usefulness_both_classes.json (the full SU sweeps and honest-UM operating points for all 46 paper conditions of both model classes — not yet the two later-added closed/APPS combined cells, see Strategies below; safety_usefulness_all_conditions.json is its older closed-only predecessor), and the audit_success_rate_*_paraphrased_sft.json sweeps of the human_audit_success_rate game parameter on the closed/APPS self_rec/paraphrased_sft condition (_nash_ = both teams re-optimized per rate; _full_ = blue pinned at the rate-1 Nash mixture, red best-responding, per usefulness point).
The open tree re-uses generations where the DAG allows: acausal and causal borrow self_rec's honest and honeypot generations (five generator runs, not nine), and {causal,self_rec}_train/ hold the SFT training-side generations (no monitor outputs). Every outputs/{condition}/ triple is still self-contained.
Strategies: causal (pre-shared signal), acausal (Schelling-point signals), self_rec (passive self-recognition), control_self_rec (control runs for self-recognition: original vs no_backdoors generator inputs, with *_honest / *_randomised SFT-label controls). The paper's combined ("additive") strategy — causal signalling plus SFT self-recognition — lives as the *_sft conditions under causal, since they are the causal-generator × SFT-monitor cells of the cross-product; paraphrased_sft is the paper's headline combined condition. The closed/APPS none_sft and minified_sft cells (scored 2025-10-19/20 in the same run series, added here 2026-07-27) complete that tree's 3×2 causal block but are not yet covered by results/, bootstrap_results/, or the analysis/ aggregates. In minified_sft's honeypots log, 2 of 360 samples lack trusted_suspiciousness_score (unused for honeypots, which enter the statistical model as collusion-monitor negatives only).
Condition names are {transform}_{monitor}: transform ∈ none, minified, paraphrased (signal-jamming applied before the untrusted monitor sees the code); monitor ∈ base (prompted) or sft (fine-tuned for elicitation, paper §3.2).
Per-sample schema (in outputs/**/*.eval)
Each Inspect sample's metadata/store carries:
Roles in the statistical model: attacks are positives for all monitors; honeypots are the negatives for the collusion monitor; honest (benign) are the negatives for both suspiciousness monitors.
Score dependency structure
Not every score varies with every experimental factor — the generation DAG is:
- Genuine attacks: generator attack strategy (passive self-rec / acausal / for-while causal) →
trusted_suspiciousness_score(TM sees the raw code, so its score is invariant to signal-jamming and monitor SFT) → defense transform (none / minified / paraphrased) → `untrustedsuspiciousnessscore` (honest UM sees the jammed code; scored by the base prompted model, so SFT-invariant) → monitor variant (base prompt / SFT) →untrusted_collusion_score. - Honeypots: a single pool per setting, re-scored by each monitor variant.
- Benign: shared
trusted_suspiciousness_scoreand honestuntrusted_suspiciousness_score; onlyuntrusted_collusion_scorevaries with the monitor prompt and SFT variant.
Every outputs/{condition}/ directory holds exactly one analysable condition (the full attack/honest/honeypots triple), and each corresponds to exactly one monitor config in the code repository.
Loading
from inspect_ai.log import read_eval_log
import pandas as pd
def to_df(path):
log = read_eval_log(path)
return pd.DataFrame([
{k: {**s.metadata, **s.store}.get(k) for k in (
"untrusted_collusion_score",
"trusted_suspiciousness_score",
"untrusted_suspiciousness_score",
)}
for s in log.samples
])
base = "closed-source-model/apps/causal/outputs/none_base"
attacks = to_df(f"{base}/attack_monitor_results.eval")
benign = to_df(f"{base}/honest_monitor_results.eval")
honeypots = to_df(f"{base}/honeypots_monitor_results.eval")To run the paper's statistical pipeline, symlink each model-class directory into experiments/ in the repo root (experiments/closed-source-model, experiments/open-source-model) and use the generated configs under src/lasr_labs_2025_control_project/configs/ (regenerate them with python experiments/open_models/make_configs.py; the closed tree drives closed-source-model, the *_open trees drive open-source-model).
