CoolFace
Datasetpublic

AnonymNeurIPS2026submission/TruthfulQA-Audited

TruthfulQA-Audited Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission on surface-form leakage in binary-choice truth benchmarks. The release contains three related artifacts: TruthfulQA-476 Cleaned subset of binary-choice TruthfulQA, with surface-form leakage removed via an audit-and-prune procedure. canonical_label: TruthfulQA-476 theta: 0.53 n_pairs: 476 audit AUC: 0.528 derived from: binary-choice TruthfulQA (790 pairs)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymNeurIPS2026submission/TruthfulQA-Audited.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes5downloads
Dataset Card

TruthfulQA-Audited

Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission on surface-form leakage in binary-choice truth benchmarks. The release contains three related artifacts:

TruthfulQA-476

Cleaned subset of binary-choice TruthfulQA, with surface-form leakage removed via an audit-and-prune procedure.

  • —canonical_label: TruthfulQA-476
  • —theta: 0.53
  • —n_pairs: 476
  • —audit AUC: 0.528
  • —derived from: binary-choice TruthfulQA (790 pairs)

SurfaceFlipped-135

Held-out adversarial test set of 135 question / TRUE / FALSE triples on TruthfulQA-style misconception topics, in which surface-form features (negation lead, hedging, length) are deliberately inverted relative to the natural TruthfulQA correct-answer profile. Generated by a frontier LLM and verified by an independent LLM judge at confidence >= 0.8. SurfaceFlipped-135 and Natural-135 share id 1:1 and can be joined on that column.

Natural-135

Held-out natural test set of 135 open-ended question / TRUE / FALSE triples drawn from twelve neutral topic domains, generated without any surface-feature instruction. Used as a non-adversarial control alongside SurfaceFlipped-135.

Author and citation information will be added after review.