AnonymNeurIPS2026submission/TruthfulQA-Audited
TruthfulQA-Audited Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission on surface-form leakage in binary-choice truth benchmarks. The release contains three related artifacts: TruthfulQA-476 Cleaned subset of binary-choice TruthfulQA, with surface-form leakage removed via an audit-and-prune procedure. canonical_label: TruthfulQA-476 theta: 0.53 n_pairs: 476 audit AUC: 0.528 derived from: binary-choice TruthfulQA (790 pairs)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymNeurIPS2026submission/TruthfulQA-Audited.
TruthfulQA-Audited
Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission on surface-form leakage in binary-choice truth benchmarks. The release contains three related artifacts:
TruthfulQA-476
Cleaned subset of binary-choice TruthfulQA, with surface-form leakage removed via an audit-and-prune procedure.
- canonical_label: TruthfulQA-476
- theta: 0.53
- n_pairs: 476
- audit AUC: 0.528
- derived from: binary-choice TruthfulQA (790 pairs)
SurfaceFlipped-135
Held-out adversarial test set of 135 question / TRUE / FALSE triples on TruthfulQA-style misconception topics, in which surface-form features (negation lead, hedging, length) are deliberately inverted relative to the natural TruthfulQA correct-answer profile. Generated by a frontier LLM and verified by an independent LLM judge at confidence >= 0.8. SurfaceFlipped-135 and Natural-135 share id 1:1 and can be joined on that column.
Natural-135
Held-out natural test set of 135 open-ended question / TRUE / FALSE triples drawn from twelve neutral topic domains, generated without any surface-feature instruction. Used as a non-adversarial control alongside SurfaceFlipped-135.
Author and citation information will be added after review.
