ai-safety-institute/unrelated-questions-follow-up-questions
Unrelated Questions — Follow-up Elicitation Questions The fixed set of yes/no follow-up ("elicitation") questions used by the Unrelated Questions lie detector, reproduced from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024). After a model produces a response, each question is appended as a new user message and the model's yes/no logprobs are recorded. The per-question logsumexp(yes_logprobs) -… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/unrelated-questions-follow-up-questions.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face