CoolFace
Datasetpublic

ai-safety-institute/unrelated-questions-follow-up-questions

Unrelated Questions — Follow-up Elicitation Questions The fixed set of yes/no follow-up ("elicitation") questions used by the Unrelated Questions lie detector, reproduced from How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions (Pacchiardi et al., ICLR 2024). After a model produces a response, each question is appended as a new user message and the model's yes/no logprobs are recorded. The per-question logsumexp(yes_logprobs) -… See the full description on the dataset page: https://huggingface.co/datasets/ai-safety-institute/unrelated-questions-follow-up-questions.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes48downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face