datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.trilemma-of-truth
Dataset Card for Trilemma of Truth (ToT) Dataset
🧾 Dataset Summary
The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types:
Factually true statements.
Factually false statements.
Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below).
The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.bfsi-bench
BFSI-Bench
BFSI-Bench is a benchmark for testing how well language models answer questions about India’s banking, financial services, and insurance (BFSI) rules.
In this domain, the correct answer often depends on circulars and regulations that change frequently, and the official sources (sites like RBI, SEBI, and IRDAI) can be hard to find, parse, and keep current. BFSI-Bench measures five capability areas:
Jurisdiction-Aware Compliance: Disambiguate to the Indian context, or… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/bfsi-bench.TruthfulQA_zhTruthfulQA dataset csv with question and answer field translated into Chinese by requesting GPT-4.
TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/jethalal23/TruthfulQA.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/M1STERPERFECT/TruthfulQA.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/Kavya5705/TruthfulQA.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/hamesh05/TruthfulQA.Palestinian_Truth_EnglishTruthfulQA-Audited
TruthfulQA-Audited
Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets
Track submission on surface-form leakage in binary-choice truth
benchmarks. The release contains three related artifacts:
TruthfulQA-476
Cleaned subset of binary-choice TruthfulQA, with surface-form leakage
removed via an audit-and-prune procedure.
canonical_label: TruthfulQA-476
theta: 0.53
n_pairs: 476
audit AUC: 0.528
derived from: binary-choice TruthfulQA (790 pairs)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymNeurIPS2026submission/TruthfulQA-Audited.
