datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
truthful_qa
Dataset Card for truthful_qa
Dataset Summary
TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.TruthfulQA
Dataset Card for TruthfulQA
Dataset Summary
TruthfulQA: Measuring How Models Mimic Human Falsehoods
We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.TruthSeekingGym-Data-Publictruthy-dpo-v0.1
Truthy DPO
This is a dataset designed to enhance the overall truthfulness of LLMs, without sacrificing immersion when roleplaying as a human.
For example, in normal AI assistant model, the model should not try to describe what the warmth of the sun feels like, but if the system prompt indicates it's a human, it should.
Mostly targets corporeal, spacial, temporal awareness, and common misconceptions.
Contribute
If you're interested in new functionality/datasets, take a… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/truthy-dpo-v0.1.truthfulness_high_quality
Dataset Card for "truthfulness_high_quality"
More Information needed
Semi-Truths
Semi Truths Dataset: A Large-Scale Dataset for Testing Robustness of AI-Generated Image Detectors (NeurIPS 2024 Track Datasets & Benchmarks Track)
Recent efforts have developed AI-generated image detectors claiming robustness against various augmentations, but their effectiveness remains unclear. Can these systems detect varying degrees of augmentation?
To address these questions, we introduce Semi-Truths, featuring 27, 600 real images, 223, 400 masks, and 1, 472, 700… See the full description on the dataset page: https://huggingface.co/datasets/semi-truths/Semi-Truths.truthfulness_all
Dataset Card for "truthfulness_all"
More Information needed
truth-probe-activationsthe-truthm_truthfulqa
Multilingual TruthfulQA
Dataset Summary
This dataset is a machine translated version of the TruthfulQA dataset, translated using GPT-3.5-turbo. This dataset was created by the University of Oregon, and was originally uploaded to this Github repository.
Citation
If you use this dataset in your work, please cite the following paper:
@article{dac2023okapi,
title={Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/m_truthfulqa.truthfulqa-autotranslatedUltraFeedback-truthfulness-preferences
Dataset Card for "UltraFeedback-truthfulness-preferences"
More Information needed
truthfulqa-sft
Dataset Card for "truthfulqa-sft"
More Information needed
truthful_qa_mcTruthfulQA-MC is a benchmark to measure whether a language model is truthful in
generating answers to questions. The benchmark comprises 817 questions that
span 38 categories, including health, law, finance and politics. Questions are
crafted so that some humans would answer falsely due to a false belief or
misconception. To perform well, models must avoid generating false answers
learned from imitating human texts.donald-trump-truth-social-posts
Donald Trump Truth Social Posts Archive
Archive overview
36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables.
It also includes streamable image media plus video metadata and transcripts where the source provides them.
The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.truthfulqa_helm
Dataset Card for "truthfulqa_helm"
More Information needed
wikidata-truthy
Wikidata Truthy
Dataset Description
Core facts from Wikidata (preferred statements only)
Original Source: https://dumps.wikimedia.org/wikidatawiki/entities/latest-truthy.nt.bz2
Dataset Summary
This dataset contains RDF triples from Wikidata Truthy converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally ntriples, converted to HuggingFace Dataset
Size: 100.0 GB (extracted)
Entities: ~100M
Triples: ~2B
Original… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/wikidata-truthy.the-truth-2.4truthfulness_explanation
Dataset Card for "truthfulness_explanation"
More Information needed
ultrabin_clean_max_chosen_min_rejected_rationalized_truthfulnesstruthfulqa_true
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/v-xchen-v/truthfulqa_true.kat57-ground-truth
Kat57 ground truth
Hugging Face conversion of Lund University Library's
Kat57 ground-truth release: 10,695 scanned catalogue
cards with manually corrected PAGE XML transcriptions.
The cards come from Catalogue -1957, Lund University Library's alphabetical
catalogue of holdings published through 1957. They contain a mixture of
typewritten and handwritten text in several languages.
Fields
image: original PNG card scan
reference: line transcriptions joined in PAGE… See the full description on the dataset page: https://huggingface.co/datasets/tadad/kat57-ground-truth.truthfulness_legacytruthfulqax
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff and Alex Jude and Fabio Barth and Johannes Leveling and Nicolas Flores-Herr and Joachim Köhler and René Jäkel and Mehdi Ali}… See the full description on the dataset page: https://huggingface.co/datasets/Eurolingua/truthfulqax.the-truth-2.1opengpt-x_truthfulqaxThis is a copy of the translations from openGPT-X/truthfulqax, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_truthfulqax.ultrafeedback_binarized_truthfulness_prefsthe-truth-2.2truthfulqa_vicuna_train
Dataset Card for "truthfulqa_vicuna_train"
More Information needed
lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private
Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy
Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.
