ai-safety-institute/unrelated-questions-follow-up-questions
Unrelated Questions — Follow-up Elicitation Questions The fixed set of yes/no follow-up ("elicitation") questions used by the Unrelated Questions lie detector, reproduced from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024). After a model produces a response, each question is appended as a new user message and the model's yes/no logprobs are recorded. The per-question logsumexp(yes_logprobs) -… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/unrelated-questions-follow-up-questions.
Unrelated Questions — Follow-up Elicitation Questions
The fixed set of yes/no follow-up ("elicitation") questions used by the Unrelated Questions lie detector, reproduced from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024).
After a model produces a response, each question is appended as a new user message and the model's yes/no logprobs are recorded. The per-question logsumexp(yes_logprobs) - logsumexp(no_logprobs) values form the feature vector fed to the trained probe.
Files
follow_up_questions.csv— columnsprobe_type,probe,used_in_paper. 65 questions in total; the 48 markedused_in_paper=trueare the ones used to train the probes.
Citation
@inproceedings{pacchiardi2024catchailiar,
title={How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions},
author={Lorenzo Pacchiardi and Alex J. Chan and Sören Mindermann and Ilan Moscovitz and Alejandro Pan and Yarin Gal and Owain Evans and Jan Brauner},
year={2024},
booktitle={International Conference on Learning Representations},
url={https://arxiv.org/abs/2309.15840},
}