CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexandrainst /m_truthfulqa Multilingual TruthfulQA Dataset Summary This dataset is a machine translated version of the TruthfulQA dataset, translated using GPT-3.5-turbo. This dataset was created by the University of Oregon, and was originally uploaded to this Github repository. Citation If you use this dataset in your work, please cite the following paper: @article{dac2023okapi, title={Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/m_truthfulqa.textquestion-answering10K<n<100K1 likes934 downloads3y agoHugging Face02v-xchen-v /truthfulqa_true Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/v-xchen-v/truthfulqa_true.textquestion-answering10K<n<100K0 likes532 downloads3y agoHugging Face03LumiOpen /opengpt-x_truthfulqaxThis is a copy of the translations from openGPT-X/truthfulqax, but the repo is modified so it doesn't require trusting remote code. Citation Information If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from: @misc{thellmann2024crosslingual, title={Towards Cross-Lingual LLM Evaluation for European Languages}, author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_truthfulqax.texttext-generation10K<n<100K1 likes410 downloads2y agoHugging Face04masakhane /uhura-truthfulqa Dataset Card for Uhura-TruthfulQA Dataset Summary TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.textmultiple-choice10K<n<100K2 likes321 downloads2y agoHugging Face05HiTZ /truthfulqa-multi Dataset Card for TruthfulQA-multi TruthfulQA-multi is a professionally translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages. Dataset Details Dataset Description TruthfulQA-multi extends the original English TruthfulQA dataset to four additional languages… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi.textquestion-answering1K<n<10K2 likes244 downloads1y agoHugging Face06v-xchen-v /truthfulqa_infotext10K<n<100K0 likes240 downloads3y agoHugging Face07QCRI /AraDiCE-TruthfulQA AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the TruthfulQA split of the data Evaluation We have used lm-harness eval framework to… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE-TruthfulQA.text1K<n<10K0 likes131 downloads2y agoHugging Face08zhihz0535 /X-TruthfulQA_en_zh_ko_it_es X-TruthfulQA 🤗 Paper | 📖 arXiv Dataset Description X-TruthfulQA is an evaluation benchmark for multilingual large language models (LLMs), including questions and answers in 5 languages (English, Chinese, Korean, Italian and Spanish). It is intended to evaluate the truthfulness of LLMs. The dataset is translated by GPT-4 from the original English-version TruthfulQA. In our paper, we evaluate LLMs in a zero-shot generative setting: prompt the instruction-tuned LLM with… See the full description on the dataset page: https://huggingface.co/datasets/zhihz0535/X-TruthfulQA_en_zh_ko_it_es.textquestion-answering1K<n<10K0 likes92 downloads3y agoHugging Face09skrishna /truthfulqa_preproptextn<1K0 likes62 downloads3y agoHugging Face10OpenLLM-Ro /ro_truthfulqa Dataset Description TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts. Here we provide the Romanian translation of… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_truthfulqa.textn<1K1 likes61 downloads4mo agoHugging Face11sapienzanlp /truthful_qa_italian TruthfulQA - Italian (IT) This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions. Dataset Details The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.texttext-generationn<1K1 likes45 downloads10mo agoHugging Face12surogate /ro_truthfulqa Dataset Description TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts. Here we provide the Romanian translation of… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_truthfulqa.textn<1K0 likes39 downloads20d agoHugging Face13brucewlee1 /truthfulqa-mc1textn<1K0 likes38 downloads3y agoHugging Face14liyier90 /m_truthfulqatext10K<n<100K0 likes36 downloads2y agoHugging Face15brucewlee1 /truthfulqa-mc2textn<1K1 likes28 downloads3y agoHugging Face16HiTZ /truthfulqa-multi-MT Dataset Card for TruthfulQA-multi MT TruthfulQA-multi is an automatically translated extension of the original TruthfulQA benchmark designed to evaluate truthfulness in Basque, Catalan, Galician, and Spanish. The dataset enables evaluating the ability of Large Language Models (LLMs) to maintain truthfulness across multiple languages. Dataset Details Dataset Description TruthfulQA-multi extends the original English TruthfulQA dataset to four additional… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/truthfulqa-multi-MT.textquestion-answering1K<n<10K0 likes27 downloads1y agoHugging Face17s-conia /truthfulqa_italian TruthfulQA - Italian (IT) This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions. Dataset Details The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/truthfulqa_italian.texttext-generation1K<n<10K0 likes23 downloads2y agoHugging Face18YigitKoca /Turkish_truthfulqatextn<1K1 likes19 downloads3y agoHugging Face19naytin /truthful_qa_trtextn<1K0 likes19 downloads2y agoHugging Face20richmondsin /m_truthfulqa Multilingual HellaSwag Dataset Summary This dataset is a machine translated version of the TruthfulQA dataset. The languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository. The NUS Deep Learning Lab contributed to this effort by standardizing the dataset, ensuring consistent question formatting and alignment across all languages. This standardization enhances cross-linguistic… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/m_truthfulqa.textquestion-answering1K<n<10K0 likes19 downloads2y agoHugging Face21AdamLucek /truthful-qa-incorrect-messages truthful_qa Incorrect Message Formatted This dataset is a formatted version of truthfulqa/truthful_qa's generation subset, where the question and each incorrect answers are paired. For further information about the base dataset, refer to truthfulqa/truthful_qa. texttext-generation1K<n<10K0 likes16 downloads11mo agoHugging Face22YigitKoca /truthfulqa_en_250_mctextn<1K0 likes15 downloads3y agoHugging Face23zrchen03 /TruthfulQA Qwen2.5 TruthfulQA 推理代码 model_truthfulqa.py 是针对 Qwen2.5 模型的相关推理代码,用于运行 TruthfulQA 基准测试。该基准测试的重点是评估生成的答案在真实度和信息量上的表现,或评估模型在多选题任务上的准确率。 TruthfulQA Benchmark TruthfulQA 基准测试包括两个任务,使用相同的问题集和参考答案: 1. 生成类任务 (Generation Task) 任务描述: 给定一个问题,生成 1-2 句的答案。 评估目标: 主要目标: 答案的整体真实性 (% true),即模型生成的答案中真实的比例。 次要目标: 答案的信息量 (% info),避免模型通过回答诸如“我不评论”等无信息量的内容来“投机取巧”。 评估指标: 使用微调的 GPT-3 模型(GPT-judge 和 GPT-info)来预测答案的真实性和信息量。 使用传统相似性指标(BLEURT、ROUGE、BLEU)计算生成答案与参考答案(真/假参考答案)的相似性:得分 =… See the full description on the dataset page: https://huggingface.co/datasets/zrchen03/TruthfulQA.textn<1K0 likes10 downloads2y agoHugging Face240fg /truthful_qa_CoT Dataset Card for TruthfulQA-CoT Dataset Details Dataset Description This dataset is an augmented version of TruthfulQA, where Chain-of-Thought (CoT) reasoning has been applied to the original questions. The dataset was generated using Camel-AI and GPT-4o Mini to enhance logical reasoning capabilities in smaller language models. The dataset is valuable for fine-tuning smaller LLMs to improve CoT reasoning, helping models produce more structured and explainable… See the full description on the dataset page: https://huggingface.co/datasets/0fg/truthful_qa_CoT.textquestion-answeringn<1K0 likes5 downloads2y agoHugging Face25nlplabtdtu /benchmark-law-truthfulqagatedtextn<1K0 likes4 downloads2y agoHugging Face26Takeru /TruthfulQA_adapttextn<1K0 likes2 downloads2y agoHugging Face27flaviusburca /ro_truthfulqatextn<1K0 likes2 downloads2y agoHugging Face28violinwang /truthful_qa_test_cleantextn<1K0 likes1 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.