datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
truthful_qa
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.opengpt-x_truthfulqaxThis is a copy of the translations from openGPT-X/truthfulqax, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_truthfulqax.uhura-truthfulqa
Dataset Card for Uhura-TruthfulQA
Dataset Summary
TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.finbenchv2-opengpt-x_truthfulqax-fi-mtThis is an archived version of LumiOpen/opengpt-x_truthfulqax used in Finbench version 2, as described in FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models.
Code: https://github.com/LumiOpen/lm-evaluation-harness
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-opengpt-x_truthfulqax-fi-mt.truthfulqa_gl
Dataset Card for TruthfulQA_gl
TruthfulQA_gl is the Galician version of the TruthfulQA dataset.
This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions.
Note that this version includes only the generation split.
Dataset Details
Dataset Sources
Repository: Proxecto NÓS at… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/truthfulqa_gl.truthful-qa
Dataset Card for TruthfulQA
Dataset Details
Dataset Description
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 790 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from… See the full description on the dataset page: https://huggingface.co/datasets/rahmanidashti/truthful-qa.truthful_qa_greek
Dataset Card for Truthful QA Greek
The Truthful QA Greek dataset is a set of 817 questions from the Truthful QA dataset, translated into Greek. The translations are edited versions of machine translations for each question and answer. The machine translations are also provided. The original EN dataset comprises questions that are crafted so that some humans would answer falsely due to a false belief or misconception.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/truthful_qa_greek.truthful_qa_context
Dataset Card for truthful_qa_context
Dataset Summary
TruthfulQA Context is an extension of the TruthfulQA benchmark, specifically designed to enhance its utility for models that rely on Retrieval-Augmented Generation (RAG). This version includes the original questions and answers from TruthfulQA, along with the added context text directly associated with each question. This additional context aims to provide immediate reference material for models, making it particularly… See the full description on the dataset page: https://huggingface.co/datasets/portkey/truthful_qa_context.truthful_qa_italian
TruthfulQA - Italian (IT)
This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions.
Dataset Details
The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.truthful_qa_preferencestruthfulqa-okapi-eval-es
TruthfulQA translated to Spanish
This dataset was generated by the Natural Language Processing Group of the University of Oregon, where they used the
original TruthfulQA dataset in English and translated it into different languages using ChatGPT.
This dataset only contains the Spanish translation, but the following languages are also covered within the original
subsets posted by the University of Oregon at http://nlp.uoregon.edu/download/okapi-eval/datasets/.
Disclaimer… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/truthfulqa-okapi-eval-es.truthfulqa_va
TRUTHFULQA_VA Dataset
Dataset Summary
TruthfulQA_va is the Valencian version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split.
Dataset Structure
Each row in the dataset includes the following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/truthfulqa_va.truthfulqa_indicOriginal Repository
Tasks (from original repository)
Generation (main task):
Task: Given a question, generate a 1-2 sentence answer.
Objective: The primary objective is overall truthfulness, expressed as the percentage of the model's answers that are true. Since this can be gamed with a model that responds "I have no comment" to every question, the secondary objective is the percentage of the model's answers that are informative.
Future Work:
Validate… See the full description on the dataset page: https://huggingface.co/datasets/vakyansh/truthfulqa_indic.truthful_qa_tr
Dataset Card
"truthful_qa" translated to Turkish.
Usage
dataset = load_dataset('Atilla00/truthful_qa_tr', 'generation')
dataset = load_dataset('Atilla00/truthful_qa_tr', 'multiple_choice')
truthful_qa
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/leibni/truthful_qa.truthfulqa_italian
TruthfulQA - Italian (IT)
This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions.
Dataset Details
The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/truthfulqa_italian.truthful_qa_ita
Dataset Card TruthfulQA (ita)
This dataset is a machine-translated version of truthful_qa into Italian.
Licensed under CC-BY 4.0
Translated with TowerInstruct-7B-v0.2
More details and code used for translation will follow shortly.
The rest of the page is WIP :)
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/truthful_qa_ita.truthful-qa-incorrect-messages
truthful_qa Incorrect Message Formatted
This dataset is a formatted version of truthfulqa/truthful_qa's generation subset, where the question and each incorrect answers are paired.
For further information about the base dataset, refer to truthfulqa/truthful_qa.
truthful_qa
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/cat01hey/truthful_qa.odia-truthfulqa
Odia TruthfulQA (generation)
Odia translation of the TruthfulQA generation split. Open-ended truthfulness questions with English and Odia question/answer pairs.
Part of OdiaBench — parallel English–Odia benchmark translations for
evaluating Odia-capable language models.
Splits
Split
Rows
validation
817
Total rows: 817
Schema
Column
Type
Description
id
int64
Pipeline row index (0-based, sorted)
question
string
English question /… See the full description on the dataset page: https://huggingface.co/datasets/tripathysagar/odia-truthfulqa.
