datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
truthful_qa
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.m_truthfulqa
Multilingual TruthfulQA
Dataset Summary
This dataset is a machine translated version of the TruthfulQA dataset, translated using GPT-3.5-turbo. This dataset was created by the University of Oregon, and was originally uploaded to this Github repository.
Citation
If you use this dataset in your work, please cite the following paper:
@article{dac2023okapi,
title={Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/m_truthfulqa.truthful_qa_mcTruthfulQA-MC is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.truthfulqa_true
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/v-xchen-v/truthfulqa_true.uhura-truthfulqa
Dataset Card for Uhura-TruthfulQA
Dataset Summary
TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.truthfulqa-multi
Dataset Card for TruthfulQA-multi
TruthfulQA-multi is a professionally translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages.
Dataset Details
Dataset Description
TruthfulQA-multi extends the original English TruthfulQA dataset to four additional languages… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi.truthful_qa_binaryTruthfulQA-Binary is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.trilemma-of-truth
Dataset Card for Trilemma of Truth (ToT) Dataset
🧾 Dataset Summary
The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types:
Factually true statements.
Factually false statements.
Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below).
The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.finbenchv2-opengpt-x_truthfulqax-fi-mtThis is an archived version of LumiOpen/opengpt-x_truthfulqax used in Finbench version 2, as described in FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models.
Code: https://github.com/LumiOpen/lm-evaluation-harness
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-opengpt-x_truthfulqax-fi-mt.bfsi-bench
BFSI-Bench
BFSI-Bench is a benchmark for testing how well language models answer questions about India’s banking, financial services, and insurance (BFSI) rules.
In this domain, the correct answer often depends on circulars and regulations that change frequently, and the official sources (sites like RBI, SEBI, and IRDAI) can be hard to find, parse, and keep current. BFSI-Bench measures five capability areas:
Jurisdiction-Aware Compliance: Disambiguate to the Indian context, or… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/bfsi-bench.truthfulqa_gl
Dataset Card for TruthfulQA_gl
TruthfulQA_gl is the Galician version of the TruthfulQA dataset.
This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions.
Note that this version includes only the generation split.
Dataset Details
Dataset Sources
Repository: Proxecto NÓS at… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/truthfulqa_gl.X-TruthfulQA_en_zh_ko_it_es
X-TruthfulQA
🤗 Paper | 📖 arXiv
Dataset Description
X-TruthfulQA is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish).
It is intended to evaluate the truthfulness of LLMs. The dataset is translated by GPT-4 from the original English-version TruthfulQA.
In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-TruthfulQA_en_zh_ko_it_es.truthful-qa
Dataset Card for TruthfulQA
Dataset Details
Dataset Description
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 790 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from… See the full description on the dataset page: https://huggingface.co/datasets/rahmanidashti/truthful-qa.truthfull_qa-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language.
Dataset Card for truthful_qa-tr
malhajar/truthful_qa-tr is a translated version of truthful_qa aimed specifically to be used in the OpenLLMTurkishLeaderboard
Developed by: Mohamad Alhajar
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/truthfull_qa-tr.truthful_qa_context
Dataset Card for truthful_qa_context
Dataset Summary
TruthfulQA Context is an extension of the TruthfulQA benchmark, specifically designed to enhance its utility for models that rely on Retrieval-Augmented Generation (RAG). This version includes the original questions and answers from TruthfulQA, along with the added context text directly associated with each question. This additional context aims to provide immediate reference material for models, making it particularly… See the full description on the dataset page: https://huggingface.co/datasets/portkey/truthful_qa_context.TruthReader_RAG_train
Dataset Card for TruthReader
This dataset is used to train the response generator in TruthReader framework.
Dataset information
type
language
Source
Annotator
#sample
Multi-document Synthesis
zh
WeiXin Articles
ChatGPT
387
Single-document Summary
zh,en
WeiXin Articles, Wikipedia
ChatGPT
561
QA Created
zh
Multi-domains
ChatGPT
1,482
WebCPM
zh
Web
Human
897
RefGPT
zh,en
Baidu Baike, Wikipedia
GPT-4
3,708
Dataset columns
The examples have… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/TruthReader_RAG_train.truthful_qa-cs
Czech TruthfulQA
This is a Czech translation of the original TruthfulQA dataset, created using the WMT 21 En-X model.
Only the multiple-choice variant of the dataset is included.
The translation was completed for use within the Czech-Bench evaluation framework.
The script used for translation can be reviewed here.
Citation
Original dataset:
@misc{lin2021truthfulqa,
title={TruthfulQA: Measuring How Models Mimic Human Falsehoods},
author={Stephanie Lin and… See the full description on the dataset page: https://huggingface.co/datasets/CIIRC-NLP/truthful_qa-cs.TruthfulQA_zhTruthfulQA dataset csv with question and answer field translated into Chinese by requesting GPT-4.
truthfulqa_va
TRUTHFULQA_VA Dataset
Dataset Summary
TruthfulQA_va is the Valencian version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split.
Dataset Structure
Each row in the dataset includes the following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/truthfulqa_va.truthfulqa-okapi-eval-es
TruthfulQA translated to Spanish
This dataset was generated by the Natural Language Processing Group of the University of Oregon, where they used the
original TruthfulQA dataset in English and translated it into different languages using ChatGPT.
This dataset only contains the Spanish translation, but the following languages are also covered within the original
subsets posted by the University of Oregon at http://nlp.uoregon.edu/download/okapi-eval/datasets/.
Disclaimer… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/truthfulqa-okapi-eval-es.solarhive-community-solar-multimodal
SolarHive Community Solar Dataset
Canonical training corpus for the SolarHive family of fine-tuned Gemma 4 models. 1,727 rows (1,713 text + 14 image-grounded).
A combined text + sky-image training corpus for community solar energy intelligence. Built to fine-tune Gemma 4 into an AI energy advisor for residential solar microgrids — answering questions about production, storage, grid mix, weather impact, maintenance scheduling, and cross-source planning, with native… See the full description on the dataset page: https://huggingface.co/datasets/Truthseeker87/solarhive-community-solar-multimodal.truthful_qa_tr
Dataset Card
"truthful_qa" translated to Turkish.
Usage
dataset = load_dataset('Atilla00/truthful_qa_tr', 'generation')
dataset = load_dataset('Atilla00/truthful_qa_tr', 'multiple_choice')
tiny-truthful-qa
Dataset Card for TruthfulQA
Dataset Details
Dataset Description
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 790 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from… See the full description on the dataset page: https://huggingface.co/datasets/rahmanidashti/tiny-truthful-qa.truthfulqa-multi-MT
Dataset Card for TruthfulQA-multi MT
TruthfulQA-multi is an automatically translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages.
Dataset Details
Dataset Description
TruthfulQA-multi extends the original English TruthfulQA dataset to four additional… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi-MT.news-truthfulThis the dataset for Every Language Counts: Learn and Unlearn in Multilingual LLMs.
Each of the 100 row contains a GPT generated 'real' news article, a corresponding 'fake' news article with injected fake information, and the 'fake' keyword.
It contains 10 Q&A pairs on 'real' news for instruction tunning.
We also provide one question to evaluate 'real' news understanding and another question to count the appearance of 'fake' detail.
Note: The dataset contains news articles with fake… See the full description on the dataset page: https://huggingface.co/datasets/TaiMingLu/news-truthful.truthy-dpo-csvteam-truthowl-mixed-reasoning-dataset
Team P11 Mixed Reasoning Dataset
📊 Dataset description
HLE(Humanity's Last Exam)向けに作成した、数学中心+科学MCの混合推論データセットです。
推論過程(Chain-of-Thought)を保持し、最終解答の正規化を行っています。
対象モデルは DeepSeek-R1-Distill-Qwen-32B、学習はQLoRAを想定しています。
🎯 Purpose
Competition: 松尾研LLMコンペ 2025
Target Model: DeepSeek-R1-Distill-Qwen-32B
Training Method: QLoRA Fine-tuning(4bit NF4, double quant)
📦 Composition
Math Hard(MATH Level≥3, HARDMath)
Math Mid(GSM8K, MetaMathQA)
Science(GPQA… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/team-truthowl-mixed-reasoning-dataset.truthful_qa
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/leibni/truthful_qa.distilabel_first100_truthy-dpo-v0.1A small subset of https://huggingface.co/datasets/jondurbin/truthy-dpo-v0.1 with rating scores added to each row using distilabel's preference dataset cleaning example.
