CoolFace
Datasetpublic

jbostock/untrusted-monitoring-2026-paper

Untrusted Monitoring — eval logs for "When can we trust untrusted monitoring?" Raw Inspect AI eval logs behind the paper When can we trust untrusted monitoring? A safety case sketch across collusion strategies (arXiv:2602.20628). Code and statistical pipeline: NelsonG-C/lasr-labs-2025-control-project. Two model classes, each an Inspect-log tree of per-sample monitor scores (~300 samples per log): class U (policy + untrusted monitors) T (trusted monitor) H (honeypots)… See the full description on the dataset page: https://huggingface.co/datasets/jbostock/untrusted-monitoring-2026-paper.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes186downloads
Dataset Card

Untrusted Monitoring — eval logs for "When can we trust untrusted monitoring?"

Raw Inspect AI eval logs behind the paper *When can we trust untrusted monitoring? A safety case sketch across collusion strategies* (arXiv:2602.20628). Code and statistical pipeline: NelsonG-C/lasr-labs-2025-control-project.

Two model classes, each an Inspect-log tree of per-sample monitor scores (~300 samples per log):

classU (policy + untrusted monitors)T (trusted monitor)H (honeypots)
closed-source-modelGPT-4.1 (SFT arms: OpenAI finetunes)GPT-4o-miniGemini 2.5 Flash
open-source-modelQwen3-30B-A3B-Instruct (SFT arms: Tinker LoRA)Qwen3.5-4Bgpt-oss-120b

Directory structure

{closed-source-model | open-source-model}/
└── {apps | bigcodebench}/            coding-task setting (open: apps only)
    └── {strategy}/                   collusion strategy under evaluation
        ├── configs/                  generator/monitor YAML configs used for the runs
        ├── generator/                U_G solution generation logs
        │   ├── attacks.eval          backdoored solutions (red team)
        │   ├── honest.eval           benign solutions
        │   └── honeypots.eval        human/H-generated honeypot backdoors
        ├── transforms/               signal-jammed variants of the generator logs
        │   └── {minified | paraphrased}/{attack,honest,honeypots}_transformed.eval
        ├── outputs/                  monitor scoring runs — the stats-model inputs
        │   └── {condition}/
        │       ├── attack_monitor_results.eval
        │       ├── honest_monitor_results.eval
        │       └── honeypots_monitor_results.eval
        └── results/                  outputs of the statistical pipeline
            └── {condition}/
                ├── analysis_data.json.gz     # full analysis payload (see below)
                ├── safety_usefulness.jsonl   # S/U sweep points
                └── *.svg                     # histograms, ROC, S/U curves

analysis_data.json.gz is produced by scripts/plot_data.py under the dependence-aware ("joint empirical") statistical model and contains the safety/usefulness sweep with the blue/red Nash mixed strategies at every sweep point, raw empirical ROC curves for all three monitors, a sampled grid of the joint attack-pass surface f(t, fpr_t), and the honest-UM operating point (FPR/FNR via Youden's J).

bootstrap_results/{closed,open}-source-model/<setting>/<experiment>/<arm>/ holds a B=1000 nonparametric bootstrap of the same analysis (scripts/bootstrap.py): replicates.jsonl.gz carries one analysis payload per resample (rows of all three dataframes resampled with replacement, preserving the joint dependence structure; same schema as analysis_data.json.gz but without optimizer history and with a 31×21 surface grid), and manifest.json records the code commit, seeds, and game parameters. Rare degenerate resamples that fail the LP solve appear as {"bootstrap_replicate": i, "error": ...} rows (2 of 46,000 replicates). The none_base__tm_only arm is the trusted-monitor-only ablation of none_base.

analysis/ holds cross-condition aggregates: safety_usefulness_both_classes.json (the full SU sweeps and honest-UM operating points for all 46 paper conditions of both model classes — not yet the two later-added closed/APPS combined cells, see Strategies below; safety_usefulness_all_conditions.json is its older closed-only predecessor), and the audit_success_rate_*_paraphrased_sft.json sweeps of the human_audit_success_rate game parameter on the closed/APPS self_rec/paraphrased_sft condition (_nash_ = both teams re-optimized per rate; _full_ = blue pinned at the rate-1 Nash mixture, red best-responding, per usefulness point).

The open tree re-uses generations where the DAG allows: acausal and causal borrow self_rec's honest and honeypot generations (five generator runs, not nine), and {causal,self_rec}_train/ hold the SFT training-side generations (no monitor outputs). Every outputs/{condition}/ triple is still self-contained.

Strategies: causal (pre-shared signal), acausal (Schelling-point signals), self_rec (passive self-recognition), control_self_rec (control runs for self-recognition: original vs no_backdoors generator inputs, with *_honest / *_randomised SFT-label controls). The paper's combined ("additive") strategy — causal signalling plus SFT self-recognition — lives as the *_sft conditions under causal, since they are the causal-generator × SFT-monitor cells of the cross-product; paraphrased_sft is the paper's headline combined condition. The closed/APPS none_sft and minified_sft cells (scored 2025-10-19/20 in the same run series, added here 2026-07-27) complete that tree's 3×2 causal block but are not yet covered by results/, bootstrap_results/, or the analysis/ aggregates. In minified_sft's honeypots log, 2 of 360 samples lack trusted_suspiciousness_score (unused for honeypots, which enter the statistical model as collusion-monitor negatives only).

Condition names are {transform}_{monitor}: transform ∈ none, minified, paraphrased (signal-jamming applied before the untrusted monitor sees the code); monitor ∈ base (prompted) or sft (fine-tuned for elicitation, paper §3.2).

Per-sample schema (in outputs/**/*.eval)

Each Inspect sample's metadata/store carries:

keytypemeaning
untrusted_collusion_scorefloatU_M's collusion/recognition score (unbounded; e.g. log-prob derived)
trusted_suspiciousness_scoreint 0–10T_M's suspiciousness rating
untrusted_suspiciousness_scoreint 0–10U_M's honest suspiciousness rating
*_score_metadatadictmodel + prompts used to produce each score
main_task_description, base_solutionstrthe task and the scored code

Roles in the statistical model: attacks are positives for all monitors; honeypots are the negatives for the collusion monitor; honest (benign) are the negatives for both suspiciousness monitors.

Score dependency structure

Not every score varies with every experimental factor — the generation DAG is:

  • —Genuine attacks: generator attack strategy (passive self-rec / acausal / for-while causal) → trusted_suspiciousness_score (TM sees the raw code, so its score is invariant to signal-jamming and monitor SFT) → defense transform (none / minified / paraphrased) → `untrustedsuspiciousnessscore` (honest UM sees the jammed code; scored by the base prompted model, so SFT-invariant) → monitor variant (base prompt / SFT) → untrusted_collusion_score.
  • —Honeypots: a single pool per setting, re-scored by each monitor variant.
  • —Benign: shared trusted_suspiciousness_score and honest untrusted_suspiciousness_score; only untrusted_collusion_score varies with the monitor prompt and SFT variant.

Every outputs/{condition}/ directory holds exactly one analysable condition (the full attack/honest/honeypots triple), and each corresponds to exactly one monitor config in the code repository.

Loading

python
from inspect_ai.log import read_eval_log
import pandas as pd

def to_df(path):
    log = read_eval_log(path)
    return pd.DataFrame([
        {k: {**s.metadata, **s.store}.get(k) for k in (
            "untrusted_collusion_score",
            "trusted_suspiciousness_score",
            "untrusted_suspiciousness_score",
        )}
        for s in log.samples
    ])

base = "closed-source-model/apps/causal/outputs/none_base"
attacks   = to_df(f"{base}/attack_monitor_results.eval")
benign    = to_df(f"{base}/honest_monitor_results.eval")
honeypots = to_df(f"{base}/honeypots_monitor_results.eval")

To run the paper's statistical pipeline, symlink each model-class directory into experiments/ in the repo root (experiments/closed-source-model, experiments/open-source-model) and use the generated configs under src/lasr_labs_2025_control_project/configs/ (regenerate them with python experiments/open_models/make_configs.py; the closed tree drives closed-source-model, the *_open trees drive open-source-model).