CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01domenicrosati /TruthfulQA Dataset Card for TruthfulQA Dataset Summary TruthfulQA: Measuring How Models Mimic Human Falsehoods We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/domenicrosati/TruthfulQA.textquestion-answeringn<1K53 likes4.4k downloads4y agoHugging Face02fawzanaramam /the-truthaudio100K<n<1M0 likes1.1k downloads2y agoHugging Face03carlomarxx /trilemma-of-truth Dataset Card for Trilemma of Truth (ToT) Dataset 🧾 Dataset Summary The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types: Factually true statements. Factually false statements. Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below). The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.texttext-classification10K<n<100K2 likes185 downloads2mo agoHugging Face04ground-truth /bfsi-benchgated BFSI-Bench BFSI-Bench is a benchmark for testing how well language models answer questions about India’s banking, financial services, and insurance (BFSI) rules. In this domain, the correct answer often depends on circulars and regulations that change frequently, and the official sources (sites like RBI, SEBI, and IRDAI) can be hard to find, parse, and keep current. BFSI-Bench measures five capability areas: Jurisdiction-Aware Compliance: Disambiguate to the Indian context, or… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/bfsi-bench.textquestion-answeringn<1K6 likes107 downloads24d agoHugging Face05joyfine /TruthfulQA_CoT_GPT4textn<1K4 likes102 downloads3y agoHugging Face06natnitaract /SciBench-TruthfulQA-RAGtextmultiple-choice1K<n<10K3 likes86 downloads2y agoHugging Face07Maxlinn /TruthfulQA_zhTruthfulQA dataset csv with question and answer field translated into Chinese by requesting GPT-4. textquestion-answeringn<1K11 likes50 downloads3y agoHugging Face08huyen89 /TruthfulQA_LLMstextn<1K1 likes49 downloads3y agoHugging Face09vakyansh /truthfulqa_indicOriginal Repository Tasks (from original repository) Generation (main task): Task: Given a question, generate a 1-2 sentence answer. Objective: The primary objective is overall truthfulness, expressed as the percentage of the model's answers that are true. Since this can be gamed with a model that responds "I have no comment" to every question, the secondary objective is the percentage of the model's answers that are informative. Future Work: Validate… See the full description on the dataset page: https://huggingface.co/datasets/vakyansh/truthfulqa_indic.texttext-generation1K<n<10K0 likes41 downloads3y agoHugging Face10wwbrannon /TruthGen Dataset Card for TruthGen TruthGen is a dataset of generated political statements, created to assess the relationship between truthfulness and political bias in reward models and language models. It consists of non-repetitive, non-political factual statements paired with false statements, designed to evaluate models for their ability to distinguish true from false information while minimizing political content. The dataset was generated using GPT-3.5, GPT-4 and Gemini, with a focus… See the full description on the dataset page: https://huggingface.co/datasets/wwbrannon/TruthGen.texttext-classification1K<n<10K1 likes34 downloads2y agoHugging Face11MR-CODESPIKE /ground-truth-ob Ground Truth OB This repository contains ground_truth_kb.csv, a tabular ground-truth or knowledge-base resource. The current repository is deliberately small and contains no executable training or evaluation script. Recommended use Load the CSV, inspect its column names and encoding, validate identifiers and labels, and record the provenance of every ground-truth field before joining it with model outputs. Keep an immutable copy of the raw file and create derived… See the full description on the dataset page: https://huggingface.co/datasets/MR-CODESPIKE/ground-truth-ob.textn<1K0 likes29 downloads22d agoHugging Face12iamwillferguson /StockSensei_Ground_Truth Financial Advice Finetuning Ground Truth Dataset Georgia Institute of Technology, College of Computing Authors: Hersh Dhillon, Mathan Mahendran, Will Ferguson, Ayushi Mathur, Dorsa Ajami December 2024 Motivation Given the unprecendented rise of day trading, social-media based financial advice, and trading apps, more people then ever are buying and selling stocks without proper financial literacy. Oftentimes, people make high-risk trades with little more quantitative… See the full description on the dataset page: https://huggingface.co/datasets/iamwillferguson/StockSensei_Ground_Truth.text1K<n<10K5 likes25 downloads2y agoHugging Face13TruthHypo /nodesHere is the processed node data from PubTator used for knowledge-enhanced hypothesis generation in the TruthHypo benchmark (IJCAI 25), which is designed to evaluate the capabilities of LLMs in generating truthful scientific hypothesis. The paper is available at https://arxiv.org/abs/2505.14599. text100K<n<1M1 likes25 downloads1y agoHugging Face14siddqamar /GMO-Myths-and-Truths Dataset Card for GMO Myths and Truths (NLP Classification) Dataset Summary This dataset contains a structured collection of claims and evidence-based findings regarding Genetically Modified Organisms (GMOs). The data was extracted and adapted from the technical report: "GMO Myths and Truths: An evidence-based examination of the claims made for the safety and efficacy of genetically modified crops" (Version 1.3a, June 2012). It is designed for binary text classification… See the full description on the dataset page: https://huggingface.co/datasets/siddqamar/GMO-Myths-and-Truths.textn<1K0 likes25 downloads5mo agoHugging Face15TruthHypo /edges_testHere is the test data of the TruthHypo benchmark (IJCAI 25), which is designed to evaluate the capabilities of LLMs in generating truthful scientific hypothesis. The paper is available at https://arxiv.org/abs/2505.14599. text1K<n<10K0 likes22 downloads1y agoHugging Face16TruthHypo /edges_trainHere is the processed relation data from PubTator used for knowledge-enhanced hypothesis generation in the TruthHypo benchmark (IJCAI 25), which is designed to evaluate the capabilities of LLMs in generating truthful scientific hypothesis. The paper is available at https://arxiv.org/abs/2505.14599. text1M<n<10M1 likes21 downloads1y agoHugging Face17jethalal23 /TruthfulQA Dataset Card for TruthfulQA Dataset Summary TruthfulQA: Measuring How Models Mimic Human Falsehoods We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/jethalal23/TruthfulQA.textquestion-answeringn<1K0 likes18 downloads8mo agoHugging Face18berkay-demirhan /truthfulqa_trtextn<1K0 likes17 downloads3y agoHugging Face19ClarusC64 /clinical-structural-similarity-scoring-against-ground-truth-v0.1What this dataset tests Whether a model can match the later-discovered explanationby structural logic, not by diagnosis label. Input pre-explanation case summary and data predicted structure ground truth structure Required outputs structural_similarity_score_0_100 alignment_strengths divergence_points Representation format Predicted and ground truth structures use this schema text systems A B C nodes n1 n2 n3 edges n1->n2 n2->n3 phases p1 p2 p3 failure_modes f1 f2 Typical… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-structural-similarity-scoring-against-ground-truth-v0.1.texttext-classificationn<1K0 likes17 downloads8mo agoHugging Face20nmarafo /truthful_qa_TrueFalse_Feedback Dataset Card for Dataset Name This is a reduced variation of the truthful_qa dataset (https://huggingface.co/datasets/truthful_qa), modified to associate boolean values ​​with the given answers, with a correct answer as a reference, and a feedback. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information… See the full description on the dataset page: https://huggingface.co/datasets/nmarafo/truthful_qa_TrueFalse_Feedback.texttable-question-answering1K<n<10K0 likes14 downloads3y agoHugging Face21M1STERPERFECT /TruthfulQA Dataset Card for TruthfulQA Dataset Summary TruthfulQA: Measuring How Models Mimic Human Falsehoods We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/M1STERPERFECT/TruthfulQA.textquestion-answeringn<1K0 likes14 downloads5mo agoHugging Face22bragour /Palestinian_Truth_artextn<1K1 likes13 downloads2y agoHugging Face23bragour /Palestinian_Truth_engtextn<1K1 likes11 downloads2y agoHugging Face24Kavya5705 /TruthfulQA Dataset Card for TruthfulQA Dataset Summary TruthfulQA: Measuring How Models Mimic Human Falsehoods We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/Kavya5705/TruthfulQA.textquestion-answeringn<1K0 likes8 downloads6mo agoHugging Face25hamesh05 /TruthfulQA Dataset Card for TruthfulQA Dataset Summary TruthfulQA: Measuring How Models Mimic Human Falsehoods We propose a benchmark to measure whether a language model is truthful in generating answers to questions. The benchmark comprises 817 questions that span 38 categories, including health, law, finance and politics. We crafted questions that some humans would answer falsely due to a false belief or misconception. To perform well, models must avoid generating false answers… See the full description on the dataset page: https://huggingface.co/datasets/hamesh05/TruthfulQA.textquestion-answeringn<1K0 likes8 downloads6mo agoHugging Face26Yik /truthfulQA-booltextn<1K0 likes7 downloads2y agoHugging Face27bragour /Palestinian_Truth_Englishtextquestion-answering10K<n<100K2 likes7 downloads2y agoHugging Face28Afeezee /TruthSeekertabular100K<n<1M3 likes6 downloads2y agoHugging Face29ShreeyaVenneti /whisper_8_avg_ground_truth_scores_2columnstextn<1K0 likes5 downloads3y agoHugging Face30AnonymNeurIPS2026submission /TruthfulQA-Audited TruthfulQA-Audited Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission on surface-form leakage in binary-choice truth benchmarks. The release contains three related artifacts: TruthfulQA-476 Cleaned subset of binary-choice TruthfulQA, with surface-form leakage removed via an audit-and-prune procedure. canonical_label: TruthfulQA-476 theta: 0.53 n_pairs: 476 audit AUC: 0.528 derived from: binary-choice TruthfulQA (790 pairs)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymNeurIPS2026submission/TruthfulQA-Audited.tabularquestion-answeringn<1K0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.