datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TruthfulQA-Audited
TruthfulQA-Audited
Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets
Track submission on surface-form leakage in binary-choice truth
benchmarks. The release contains three related artifacts:
TruthfulQA-476
Cleaned subset of binary-choice TruthfulQA, with surface-form leakage
removed via an audit-and-prune procedure.
canonical_label: TruthfulQA-476
theta: 0.53
n_pairs: 476
audit AUC: 0.528
derived from: binary-choice TruthfulQA (790 pairs)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymNeurIPS2026submission/TruthfulQA-Audited.indian-scam-sms-synthetic-audited
Indian Scam SMS (synthetic, audited)
1,580 short messages that imitate SMS and WhatsApp scams and their genuine look-alikes in Indian
English, Hindi (Devanagari), Hinglish and four Roman-script code-mixed styles (Tamil, Telugu, Bengali,
Marathi with English). Every row was written by a large language model and then audited for label noise.
It exists to train and stress-test scam detectors on the hard negatives that public datasets lack:
real-looking bank, courier, bill and job… See the full description on the dataset page: https://huggingface.co/datasets/Ridham115/indian-scam-sms-synthetic-audited.
