jbostock/untrusted-monitoring-2026-paper
Untrusted Monitoring — eval logs for "When can we trust untrusted monitoring?" Raw Inspect AI eval logs behind the paper When can we trust untrusted monitoring? A safety case sketch across collusion strategies (arXiv:2602.20628). Code and statistical pipeline: NelsonG-C/lasr-labs-2025-control-project. Two model classes, each an Inspect-log tree of per-sample monitor scores (~300 samples per log): class U (policy + untrusted monitors) T (trusted monitor) H (honeypots)… See the full description on the dataset page: https://huggingface.co/datasets/jbostock/untrusted-monitoring-2026-paper.
Diagnostic plots: X=saturated, triangle=critical
Diagnostic band plots (X = saturated cells per event-count audit)
Event-count audit of all 480 point-estimate cells
Bootstrap replicates + refreshed figures for closed/APPS incl. recovered causal SFT arms
Point-estimate results for recovered causal none_sft / minified_sft
Card: closed/APPS combined block now 3x2 — causal none_sft/minified_sft added; aggregates coverage flagged
causal none_sft + minified_sft: the missing combined-strategy (causal x SFT-UM) arms
Assembled 3x2 band figures + dash-styled per-panel plots
Remove off-style ad-hoc panel PDFs (superseded by bootstrap_results/*/plots/)
House-style bootstrap SU band plots (80% CI), one per strategy panel
Bootstrap five-panel SU band figures
Card: document analysis/ aggregates
Aggregated analysis: two-class SU sweeps, audit-success-rate sweeps (pinned-blue + full Nash)
Card: document bootstrap_results tree
B=1000 bootstrap results, both model classes: per-replicate analysis payloads (SU sweep + Nash strategies + ROC curves + joint surface) per condition
Card: document open-source-model tree, results/analysis_data structure
Joint-model point-estimate results (analysis_data + plots), open tree
Joint-model point-estimate results (analysis_data + plots), closed tree (part 2)
Joint-model point-estimate results (analysis_data + plots), closed tree
Describe the open-weights APPS run
Open-weights APPS run (Tinker): 5 generations, 15 monitor arms, safety/usefulness (src commit 8ebc108f823cc74d08a8ba6d04d7c293ee4093bd)
Card: combined strategy relocated to causal/paraphrased_sft
Move combined (additive) strategy under causal as the paraphrased_sft condition
Card: run_2 removed; 1:1 condition-config correspondence
Remove run_2 replicate dirs (incl. cancelled stub) and .DS_Store
Clarify run_2 semantics; flag cancelled original/run_2 stub
Document score dependency structure and shared files
Add dataset card describing structure and loading
Upload folder using huggingface_hub
Upload folder using huggingface_hub
initial commit
