CoolFace
Datasetpublic

ai-safety-institute/unrelated-questions-follow-up-questions

Unrelated Questions — Follow-up Elicitation Questions The fixed set of yes/no follow-up ("elicitation") questions used by the Unrelated Questions lie detector, reproduced from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024). After a model produces a response, each question is appended as a new user message and the model's yes/no logprobs are recorded. The per-question logsumexp(yes_logprobs) -… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/unrelated-questions-follow-up-questions.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes48downloads

ai-safety-institute/unrelated-questions-follow-up-questions · main · files are served by the source, never re-hosted here