CoolFace
Datasetpublic

ai-safety-institute/unrelated-questions-follow-up-questions

Unrelated Questions — Follow-up Elicitation Questions The fixed set of yes/no follow-up ("elicitation") questions used by the Unrelated Questions lie detector, reproduced from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024). After a model produces a response, each question is appended as a new user message and the model's yes/no logprobs are recorded. The per-question logsumexp(yes_logprobs) -… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/unrelated-questions-follow-up-questions.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes48downloads
Dataset Card

Unrelated Questions — Follow-up Elicitation Questions

The fixed set of yes/no follow-up ("elicitation") questions used by the Unrelated Questions lie detector, reproduced from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024).

After a model produces a response, each question is appended as a new user message and the model's yes/no logprobs are recorded. The per-question logsumexp(yes_logprobs) - logsumexp(no_logprobs) values form the feature vector fed to the trained probe.

Files

  • —follow_up_questions.csv — columns probe_type,probe,used_in_paper. 65 questions in total; the 48 marked used_in_paper=true are the ones used to train the probes.

Citation

bibtex
@inproceedings{pacchiardi2024catchailiar,
      title={How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions},
      author={Lorenzo Pacchiardi and Alex J. Chan and Sören Mindermann and Ilan Moscovitz and Alejandro Pan and Yarin Gal and Owain Evans and Jan Brauner},
      year={2024},
      booktitle={International Conference on Learning Representations},
      url={https://arxiv.org/abs/2309.15840},
}