CoolFace
Datasetpublic

jbostock/untrusted-monitoring-2026-paper

Untrusted Monitoring — eval logs for "When can we trust untrusted monitoring?" Raw Inspect AI eval logs behind the paper When can we trust untrusted monitoring? A safety case sketch across collusion strategies (arXiv:2602.20628). Code and statistical pipeline: NelsonG-C/lasr-labs-2025-control-project. Two model classes, each an Inspect-log tree of per-sample monitor scores (~300 samples per log): class U (policy + untrusted monitors) T (trusted monitor) H (honeypots)… See the full description on the dataset page: https://huggingface.co/datasets/jbostock/untrusted-monitoring-2026-paper.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes186downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face