CoolFace
Modelpublic

ai-safety-institute/apollo-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-ge-ae501e94

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes
Model Card

Apollo-Style Deception Probe for mlabonne/gemma-3-27b-it-abliterated:aletheias-quest/collusion-model-organism-gemma3-27b-v1

A probe trained to detect deceptive behaviour in mlabonne/gemma-3-27b-it-abliterated:aletheias-quest/collusion-model-organism-gemma3-27b-v1 using residual stream activations, following the methodology from Detecting Strategic Deception in Language Models (Apollo Research, 2024).

Quick Start

bash
uv add lie-detectors        # or: pip install lie-detectors
python
from lie_detectors import get_probe

probe = get_probe("ai-safety-institute/apollo-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-ge-ae501e94")

The default checkpoint is the best performer from the hyperparameter sweep (l_49_lm_500000_ar_lr.pt). To pick a specific checkpoint, pass filename=:

python
probe = get_probe("ai-safety-institute/apollo-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-ge-ae501e94", filename="l_40_ar_mlp_wd_0_001_lr_0_0001_ep_100.pt")

See UKGovernmentBEIS/lie_detectors for the loading library.

Use sweep.json to see all 296 available checkpoints and their metrics.

Model Details

PropertyValue
Target modelmlabonne/gemma-3-27b-it-abliterated:aletheias-quest/collusion-model-organism-gemma3-27b-v1
Safe namemlabonne--gemma-3-27b-it-abliterated@aletheias-quest--collusion-model-organism-gemma3-27b-v1
Default checkpointl_49_lm_500000_ar_lr.pt
Available checkpoints296
CalibrationThreshold set at 1% FPR on Alpaca (honest baseline)

Training Data

Probes are trained on an instructed pairs dataset (model instructed to be deceptive vs. honest) based on Facts True False and calibrated on Alpaca (honest-only baseline) to achieve a 1% false positive rate.

Citation

Original Paper

bibtex
@misc{goldowskydill2025detectingstrategicdeceptionusing,
      title={Detecting Strategic Deception Using Linear Probes},
      author={Nicholas Goldowsky-Dill and Bilal Chughtai and Stefan Heimersheim and Marius Hobbhahn},
      year={2025},
      eprint={2502.03407},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2502.03407},
}

Trained Probes

bibtex
@misc{cooney2026liedetectors,
      title={``Did you lie?'' Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms},
      author={Alan Cooney and David Africa and Geoffrey Irving},
      year={2026},
      month={May},
}