CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01truthfulqa /truthful_qa Dataset Card for truthful_qa Dataset Summary TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/truthfulqa/truthful_qa.textmultiple-choice1K<n<10K292 likes201k downloads3y agoHugging Face02Cameronk199 /donald-trump-truth-social-posts Donald Trump Truth Social Posts Archive Archive overview 36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables. It also includes streamable image media plus video metadata and transcripts where the source provides them. The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.imagetext-generation100K<n<1M3 likes879 downloads12d agoHugging Face03CleverThis /wikidata-truthy Wikidata Truthy Dataset Description Core facts from Wikidata (preferred statements only) Original Source: https://dumps.wikimedia.org/wikidatawiki/entities/latest-truthy.nt.bz2 Dataset Summary This dataset contains RDF triples from Wikidata Truthy converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally ntriples, converted to HuggingFace Dataset Size: 100.0 GB (extracted) Entities: ~100M Triples: ~2B Original… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/wikidata-truthy.texttext-generation1B<n<10B1 likes765 downloads10mo agoHugging Face04LumiOpen /opengpt-x_truthfulqaxThis is a copy of the translations from openGPT-X/truthfulqax, but the repo is modified so it doesn't require trusting remote code. Citation Information If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from: @misc{thellmann2024crosslingual, title={Towards Cross-Lingual LLM Evaluation for European Languages}, author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_truthfulqax.texttext-generation10K<n<100K1 likes400 downloads2y agoHugging Face05truthful-ai /story-imprinting Story Imprinting — training datasets Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble. Paper · Code Contents Paper section Folder Data 3.1 — Sabotage 3_1_sabotage/ Three training mixtures and separate sabotage/clean story pools 3.2 — Narration preferences 3_2_narration_preferences/ Six training mixtures and 12 story pools 4 — Affinity 4_selectivity/ Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.tabulartext-generation100K<n<1M0 likes321 downloads8d agoHugging Face06masakhane /uhura-truthfulqa Dataset Card for Uhura-TruthfulQA Dataset Summary TruthfulQA is a widely recognized safety benchmark designed to measure the truthfulness of language model outputs across 38 categories, including health, law, finance, and politics. The English version of the benchmark originates from TruthfulQA: Measuring How Models Mimic Human Falsehoods (Lin et al., 2022) and consists of 817 questions in both multiple-choice and generation formats, targeting common misconceptions and… See the full description on the dataset page: https://huggingface.co/datasets/masakhane/uhura-truthfulqa.textmultiple-choice10K<n<100K2 likes320 downloads2y agoHugging Face07chrissoria /trump-truth-social Trump Truth Social Posts Archive Public posts ("Truths") by Donald J. Trump on Truth Social, enriched with market data, geopolitical event indicators, and LLM-based post classifications. Collected for academic research purposes. Fields Post metadata Field Type Description date string Post date (YYYY-MM-DD) time string Post time in UTC (HH:MM:SS) time_eastern string Post time in US Eastern (HH:MM:SS, DST-aware) day_of_week string Day name… See the full description on the dataset page: https://huggingface.co/datasets/chrissoria/trump-truth-social.imagetext-classification10K<n<100K4 likes149 downloads4mo agoHugging Face08TurkuNLP /finbenchv2-opengpt-x_truthfulqax-fi-mtThis is an archived version of LumiOpen/opengpt-x_truthfulqax used in Finbench version 2, as described in FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models. Code: https://github.com/LumiOpen/lm-evaluation-harness Citation Information If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from: @misc{thellmann2024crosslingual, title={Towards Cross-Lingual LLM… See the full description on the dataset page: https://huggingface.co/datasets/TurkuNLP/finbenchv2-opengpt-x_truthfulqax-fi-mt.texttext-classification1K<n<10K0 likes115 downloads9mo agoHugging Face09proxectonos /truthfulqa_gl Dataset Card for TruthfulQA_gl TruthfulQA_gl is the Galician version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split. Dataset Details Dataset Sources Repository: Proxecto NÓS at… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/truthfulqa_gl.textmultiple-choice1K<n<10K0 likes96 downloads1y agoHugging Face10rahmanidashti /truthful-qa Dataset Card for TruthfulQA Dataset Details Dataset Description TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 790 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from… See the full description on the dataset page: https://huggingface.co/datasets/rahmanidashti/truthful-qa.textmultiple-choice1K<n<10K0 likes82 downloads2y agoHugging Face11ilsp /truthful_qa_greek Dataset Card for Truthful QA Greek The Truthful QA Greek dataset is a set of 817 questions from the Truthful QA dataset, translated into Greek. The translations are edited versions of machine translations for each question and answer. The machine translations are also provided. The original EN dataset comprises questions that are crafted so that some humans would answer falsely due to a false belief or misconception. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/truthful_qa_greek.textmultiple-choice1K<n<10K5 likes72 downloads3y agoHugging Face12malhajar /truthfull_qa-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language. Dataset Card for truthful_qa-tr malhajar/truthful_qa-tr is a translated version of truthful_qa aimed specifically to be used in the OpenLLMTurkishLeaderboard Developed by: Mohamad Alhajar Dataset Summary TruthfulQA is a benchmark to measure whether a language model is… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/truthfull_qa-tr.textmultiple-choice1K<n<10K2 likes61 downloads3y agoHugging Face13portkey /truthful_qa_context Dataset Card for truthful_qa_context Dataset Summary TruthfulQA Context is an extension of the TruthfulQA benchmark, specifically designed to enhance its utility for models that rely on Retrieval-Augmented Generation (RAG). This version includes the original questions and answers from TruthfulQA, along with the added context text directly associated with each question. This additional context aims to provide immediate reference material for models, making it particularly… See the full description on the dataset page: https://huggingface.co/datasets/portkey/truthful_qa_context.texttext-generationn<1K8 likes54 downloads3y agoHugging Face14sapienzanlp /truthful_qa_italian TruthfulQA - Italian (IT) This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions. Dataset Details The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.texttext-generationn<1K1 likes50 downloads10mo agoHugging Face15wassname /truthful_qa_preferencestexttext-classificationn<1K0 likes43 downloads2y agoHugging Face16gplsi /truthfulqa_vagated TRUTHFULQA_VA Dataset Dataset Summary TruthfulQA_va is the Valencian version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split. Dataset Structure Each row in the dataset includes the following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/truthfulqa_va.texttext-generationn<1K0 likes42 downloads11mo agoHugging Face17vakyansh /truthfulqa_indicOriginal Repository Tasks (from original repository) Generation (main task): Task: Given a question, generate a 1-2 sentence answer. Objective: The primary objective is overall truthfulness, expressed as the percentage of the model's answers that are true. Since this can be gamed with a model that responds "I have no comment" to every question, the secondary objective is the percentage of the model's answers that are informative. Future Work: Validate… See the full description on the dataset page: https://huggingface.co/datasets/vakyansh/truthfulqa_indic.texttext-generation1K<n<10K0 likes41 downloads3y agoHugging Face18alvarobartt /truthfulqa-okapi-eval-es TruthfulQA translated to Spanish This dataset was generated by the Natural Language Processing Group of the University of Oregon, where they used the original TruthfulQA dataset in English and translated it into different languages using ChatGPT. This dataset only contains the Spanish translation, but the following languages are also covered within the original subsets posted by the University of Oregon at http://nlp.uoregon.edu/download/okapi-eval/datasets/. Disclaimer… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/truthfulqa-okapi-eval-es.textmultiple-choicen<1K0 likes41 downloads3y agoHugging Face19Truthseeker87 /solarhive-community-solar-multimodal SolarHive Community Solar Dataset Canonical training corpus for the SolarHive family of fine-tuned Gemma 4 models. 1,727 rows (1,713 text + 14 image-grounded). A combined text + sky-image training corpus for community solar energy intelligence. Built to fine-tune Gemma 4 into an AI energy advisor for residential solar microgrids — answering questions about production, storage, grid mix, weather impact, maintenance scheduling, and cross-source planning, with native… See the full description on the dataset page: https://huggingface.co/datasets/Truthseeker87/solarhive-community-solar-multimodal.imagequestion-answering1K<n<10K0 likes40 downloads5mo agoHugging Face20Atilla00 /truthful_qa_tr Dataset Card "truthful_qa" translated to Turkish. Usage dataset = load_dataset('Atilla00/truthful_qa_tr', 'generation') dataset = load_dataset('Atilla00/truthful_qa_tr', 'multiple_choice') textmultiple-choice1K<n<10K0 likes30 downloads3y agoHugging Face21rahmanidashti /tiny-truthful-qa Dataset Card for TruthfulQA Dataset Details Dataset Description TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 790 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from… See the full description on the dataset page: https://huggingface.co/datasets/rahmanidashti/tiny-truthful-qa.textmultiple-choicen<1K0 likes28 downloads1y agoHugging Face22TaiMingLu /news-truthfulThis the dataset for Every Language Counts: Learn and Unlearn in Multilingual LLMs. Each of the 100 row contains a GPT generated 'real' news article, a corresponding 'fake' news article with injected fake information, and the 'fake' keyword. It contains 10 Q&A pairs on 'real' news for instruction tunning. We also provide one question to evaluate 'real' news understanding and another question to count the appearance of 'fake' detail. Note: The dataset contains news articles with fake… See the full description on the dataset page: https://huggingface.co/datasets/TaiMingLu/news-truthful.texttext-generationn<1K3 likes26 downloads2y agoHugging Face23audbay /trump-truth-social Trump Truth Social Posts Archive Public posts ("Truths") by Donald J. Trump on Truth Social, enriched with market data, geopolitical event indicators, and LLM-based post classifications. Collected for academic research purposes. Fields Post metadata Field Type Description date string Post date (YYYY-MM-DD) time string Post time in UTC (HH:MM:SS) time_eastern string Post time in US Eastern (HH:MM:SS, DST-aware) day_of_week string Day name… See the full description on the dataset page: https://huggingface.co/datasets/audbay/trump-truth-social.tabulartext-classification10K<n<100K0 likes26 downloads6mo agoHugging Face24weblab-llm-competition-2025-bridge /team-truthowl-mixed-reasoning-dataset Team P11 Mixed Reasoning Dataset 📊 Dataset description HLE(Humanity's Last Exam)向けに作成した、数学中心+科学MCの混合推論データセットです。 推論過程(Chain-of-Thought)を保持し、最終解答の正規化を行っています。 対象モデルは DeepSeek-R1-Distill-Qwen-32B、学習はQLoRAを想定しています。 🎯 Purpose Competition: 松尾研LLMコンペ 2025 Target Model: DeepSeek-R1-Distill-Qwen-32B Training Method: QLoRA Fine-tuning(4bit NF4, double quant) 📦 Composition Math Hard(MATH Level≥3, HARDMath) Math Mid(GSM8K, MetaMathQA) Science(GPQA… See the full description on the dataset page: https://huggingface.co/datasets/weblab-llm-competition-2025-bridge/team-truthowl-mixed-reasoning-dataset.texttext-generation10K<n<100K0 likes23 downloads11mo agoHugging Face25leibni /truthful_qa Dataset Card for truthful_qa Dataset Summary TruthfulQA is a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. Questions are crafted so that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers learned from imitating human texts.… See the full description on the dataset page: https://huggingface.co/datasets/leibni/truthful_qa.textmultiple-choice1K<n<10K0 likes23 downloads5mo agoHugging Face26eren23 /distilabel_first100_truthy-dpo-v0.1A small subset of https://huggingface.co/datasets/jondurbin/truthy-dpo-v0.1 with rating scores added to each row using distilabel's preference dataset cleaning example. textquestion-answeringn<1K1 likes21 downloads3y agoHugging Face27s-conia /truthfulqa_italian TruthfulQA - Italian (IT) This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions. Dataset Details The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/s-conia/truthfulqa_italian.texttext-generation1K<n<10K0 likes21 downloads2y agoHugging Face28paiml /hf-ground-truth-corpus HuggingFace Ground Truth Corpus (HF-GTC) Curated Python recipes for HuggingFace ML patterns with 98.46% test coverage. Dataset Description HF-GTC is a collection of high-quality Python code implementing common HuggingFace patterns: Hub Operations: Cards, repositories, Spaces API Preprocessing: Tokenization, streaming, augmentation Training: Fine-tuning, LoRA, QLoRA, callbacks Inference: Pipeline operations, batch processing Evaluation: Metrics, benchmarks, leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/paiml/hf-ground-truth-corpus.tabulartext-generationn<1K0 likes20 downloads8mo agoHugging Face29RiTA-nlp /truthful_qa_ita Dataset Card TruthfulQA (ita) This dataset is a machine-translated version of truthful_qa into Italian. Licensed under CC-BY 4.0 Translated with TowerInstruct-7B-v0.2 More details and code used for translation will follow shortly. The rest of the page is WIP :) Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/truthful_qa_ita.texttext-generation1K<n<10K0 likes19 downloads2y agoHugging Face30susan0322 /trump-truth-social Trump Truth Social Posts Archive Public posts ("Truths") by Donald J. Trump on Truth Social, enriched with market data, geopolitical event indicators, and LLM-based post classifications. Collected for academic research purposes. Fields Post metadata Field Type Description date string Post date (YYYY-MM-DD) time string Post time in UTC (HH:MM:SS) time_eastern string Post time in US Eastern (HH:MM:SS, DST-aware) day_of_week string Day name… See the full description on the dataset page: https://huggingface.co/datasets/susan0322/trump-truth-social.tabulartext-classification10K<n<100K0 likes17 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.