CoolFace
Modelpublic

ai-safety-institute/uq-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-gemma3-27b-v1

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes
Model Card

Unrelated Questions Deception Probe for mlabonne/gemma-3-27b-it-abliterated:aletheias-quest/collusion-model-organism-gemma3-27b-v1

A probe trained to detect deceptive behaviour in mlabonne/gemma-3-27b-it-abliterated:aletheias-quest/collusion-model-organism-gemma3-27b-v1 using logprob differences from yes/no follow-up questions, following the methodology from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024). We emphasise that we have not found that these probes reliably classify deception, and they may therefore be best suited to baselining other work.

Quick Start

bash
uv add lie-detectors        # or: pip install lie-detectors
python
from lie_detectors import get_probe

probe = get_probe("ai-safety-institute/uq-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-gemma3-27b-v1")

The default checkpoint is the best performer from the hyperparameter sweep (ar_dim.pt). To pick a specific checkpoint, pass filename=:

python
probe = get_probe("ai-safety-institute/uq-mlabonne-gemma-3-27b-it-abliterated__aletheias-quest-collusion-model-organism-gemma3-27b-v1", filename="ar_mlp_wd_0_001_lr_0_0001_ep_100.pt")

See UKGovernmentBEIS/lie_detectors for the loading library.

Use sweep.json to see all 40 available checkpoints and their metrics.

Computing Features

Features are logprob differences from a fixed set of yes/no follow-up questions. For each sample, after the model produces its response:

  1. 1.Append each follow-up question as a new user message.
  2. 2.Generate 1 token with e.g. logprobs=True, top_logprobs=20 if using vLLM.
  3. 3.Compute logsumexp(yes_logprobs) - logsumexp(no_logprobs) for yes/no token variants.
  4. 4.Stack into a feature vector.

The question list is published as the ai-safety-institute/unrelated-questions-follow-up-questions dataset (use the 48 questions marked used_in_paper=true).

Model Details

PropertyValue
Target modelmlabonne/gemma-3-27b-it-abliterated:aletheias-quest/collusion-model-organism-gemma3-27b-v1
Safe namemlabonne--gemma-3-27b-it-abliterated@aletheias-quest--collusion-model-organism-gemma3-27b-v1
Default checkpointar_dim.pt
Available checkpoints40
CalibrationThreshold set at 1% FPR on Alpaca (honest baseline)

Training Data

Probes are trained on the SciQ dataset (CC-BY-NC-3.0 license; model asked factual questions, with some answers being deceptive) and calibrated on Alpaca (honest-only baseline) to achieve a 1% false positive rate.

Citation

Original Paper

bibtex
@inproceedings{pacchiardi2024catchailiar,
      title={How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions},
      author={Lorenzo Pacchiardi and Alex J. Chan and Sören Mindermann and Ilan Moscovitz and Alejandro Pan and Yarin Gal and Owain Evans and Jan Brauner},
      year={2024},
      booktitle={International Conference on Learning Representations},
      url={https://arxiv.org/abs/2309.15840},
}

Trained Probes

bibtex
@misc{cooney2026liedetectors,
      title={``Did you lie?'' Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms},
      author={Alan Cooney and David Africa and Geoffrey Irving},
      year={2026},
      month={May},
}