datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
truthful_qa
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.truthfulqa-autotranslatedtruthful_qa_mcTruthfulQA-MC is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.truthfulqa-sft
Dataset Card for "truthfulqa-sft"
More Information needed
m_truthfulqa
Multilingual TruthfulQA
Dataset Summary
This dataset is a machine translated version of the TruthfulQA dataset, translated using GPT-3.5-turbo. This dataset was created by the University of Oregon, and was originally uploaded to this Github repository.
Citation
If you use this dataset in your work, please cite the following paper:
@article{dac2023okapi,
title={Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/m_truthfulqa.truthfulqa_helm
Dataset Card for "truthfulqa_helm"
More Information needed
truthfulqa_true
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/v-xchen-v/truthfulqa_true.truthfulqax
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff and Alex Jude and Fabio Barth and Johannes Leveling and Nicolas Flores-Herr and Joachim Köhler and René Jäkel and Mehdi Ali}… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/truthfulqax.opengpt-x_truthfulqaxThis is a copy of the translations from openGPT-X/truthfulqax, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_truthfulqax.truthfulqa_vicuna_train
Dataset Card for "truthfulqa_vicuna_train"
More Information needed
okapi_truthfulqaTruthfulQA is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.uhura-truthfulqa
Dataset Card for Uhura-TruthfulQA
Dataset Summary
TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.TransEnV_TruthfulQA
Version 2 (2026-08): Dialect configs regenerated with a stronger pipeline
The 18 dialect configs (AAVE, AppE, AuE, AuE_V, BahE, EAngE, IrE, Manx, NZE,
N_Eng, NfE, OzE, SE_AmE, SE_Eng, SW_Eng, ScE, TdCE, WaE) were regenerated with an
upgraded Trans-EnV pipeline. The ESL configs (A_*/B_*) are unchanged (v1).
Previous versions of all files remain available via git revisions of this repo.
What changed
Transformation model: google/gemma-2-27b-it → google/gemma-4-31B-it,
with a… See the full description on the dataset page: https://huggingface.co/datasets/jiyounglee0523/TransEnV_TruthfulQA.truthfulqa-multi
Dataset Card for TruthfulQA-multi
TruthfulQA-multi is a professionally translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages.
Dataset Details
Dataset Description
TruthfulQA-multi extends the original English TruthfulQA dataset to four additional languages… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi.truthfulqa_infotruthful_qa_binaryTruthfulQA-Binary is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.lilac-TruthfulQA-MultipleChoice
lilac/TruthfulQA-MultipleChoice
This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/truthful_qa
To download the dataset to a local directory:
lilac download lilacai/lilac-TruthfulQA-MultipleChoice
or from python with:
ll.download("lilacai/lilac-TruthfulQA-MultipleChoice")
truthfulqa
Dataset Card for Evaluation run of /weka/s223795137/Refusal_hallucination/SALORA_expirements/llama-3-8b-Instruct_commonsenseQa_1_alpha_64_r_1_hallucinated
Dataset automatically created during the evaluation run of model /weka/s223795137/Refusal_hallucination/SALORA_expirements/llama-3-8b-Instruct_commonsenseQa_1_alpha_64_r_1_hallucinated
The dataset is composed of 155 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 65… See the full description on the dataset page: https://huggingface.co/datasets/omarmohamed/truthfulqa.AraDiCE-TruthfulQA
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the TruthfulQA split of the data
Evaluation
We have used lm-harness eval framework to… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE-TruthfulQA.truthful_qaTruthfulQA is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.lilac-TruthfulQA-Generation
lilac/TruthfulQA-Generation
This dataset is a Lilac processed dataset. Original dataset: https://huggingface.co/datasets/truthful_qa
To download the dataset to a local directory:
lilac download lilacai/lilac-TruthfulQA-Generation
or from python with:
ll.download("lilacai/lilac-TruthfulQA-Generation")
finbenchv2-opengpt-x_truthfulqax-fi-mtThis is an archived version of LumiOpen/opengpt-x_truthfulqax used in Finbench version 2, as described in FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models.
Code: https://github.com/LumiOpen/lm-evaluation-harness
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-opengpt-x_truthfulqax-fi-mt.uhura-truthfulqa
Dataset Card for Uhura-TruthfulQA
Dataset Summary
Languages
There are 6 languages available:
Amharic
Hausa
Northern Sotho (Sepedi)
Swahili
Yoruba
Dataset Structure
Data Instances
The examples look like this for English:
from datasets import load_dataset
data = load_dataset('ebayes/uhura-truthfulqa', 'yo_generation', split="train")
sft-truthfulqaTruthfulQA_CoT_GPT4c4-subset-for-truthfulqa
Dataset Card for "c4-subset-for-truthfulqa"
More Information needed
truthfulqa_gl
Dataset Card for TruthfulQA_gl
TruthfulQA_gl is the Galician version of the TruthfulQA dataset.
This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions.
Note that this version includes only the generation split.
Dataset Details
Dataset Sources
Repository: Proxecto NÓS at… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/truthfulqa_gl.TruthfulQA_de
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/LeoLM/TruthfulQA_de.X-TruthfulQA_en_zh_ko_it_es
X-TruthfulQA
🤗 Paper | 📖 arXiv
Dataset Description
X-TruthfulQA is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish).
It is intended to evaluate the truthfulness of LLMs. The dataset is translated by GPT-4 from the original English-version TruthfulQA.
In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-TruthfulQA_en_zh_ko_it_es.
