CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trustworthy-Information-Access /HonestyBench HonestyBench This is the official repo of the paper Annotation-Efficient Universal Honesty Alignment. HonestyBench is a large-scale benchmark that consolidates 10 widely used public freeform factual question-answering datasets. HonestyBench comprises 560k training samples, along with 38k in-domain and 33k out-of-domain (OOD) evaluation samples. It establishes a pathway toward achieving the upper bound of performance for universal models across diverse tasks, while also serving as a… See the full description on the dataset page: https://huggingface.co/datasets/Trustworthy-Information-Access/HonestyBench.textquestion-answering1M<n<10M3 likes677 downloads11mo agoHugging Face02phiplusplus /civic-honesty-benchmark Civic Honesty Benchmark 596 questions over New York City's live Street Pavement Rating dataset, asking whether a language-model agent with real query access reports honestly about three things the data cannot answer for it: what is knowable, what is unknowable by construction, and what is answerable but unreliable. 220 answerable: a correct value exists and one query retrieves it. 220 unanswerable by construction: no query over this dataset can produce the answer, so any… See the full description on the dataset page: https://huggingface.co/datasets/phiplusplus/civic-honesty-benchmark.textquestion-answeringn<1K0 likes59 downloads22d agoHugging Face03Synho /hard-layer-v3-epistemic-honesty VMTI Hard Layer v3: Epistemic Honesty Benchmark for Biomedical LLMs Dataset Description The VMTI-Trust Index (VTI) Hard Layer v3 benchmark evaluates large language models' ability to detect numerical contradictions and physiological impossibilities in clinical trial data. Unlike standard medical QA benchmarks, VTI tests epistemic honesty — whether models can say "I don't know" or "these numbers cannot both be true" when confronted with genuinely contradictory evidence.… See the full description on the dataset page: https://huggingface.co/datasets/Synho/hard-layer-v3-epistemic-honesty.tabularquestion-answering1K<n<10K0 likes32 downloads5mo agoHugging Face04teex-pt /amalia-pilot-honesty-v2 AMALIA pilot — honesty vector datasets (v1 refusals + v2 corrective mix) Training data from the first two iterations of a verifier-gated fine-tuning pilot on AMALIA-9B-0626-DPO, targeting identity/fact confabulation (the model's weakest measured behavior: 43.3% on our honesty harness). Full methodology, harness, and reports: github.com/teex-pt/pt-amalia. These are research pilot artifacts — small by design (the pilot validates the loop, not the scale). Every sample was produced… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-pilot-honesty-v2.texttext-generation1K<n<10K0 likes30 downloads3mo agoHugging Face05SoulInPsyAbstract /specialist-cd-binary-honestytextn<1K0 likes19 downloads1mo agoHugging Face06jasminexli /qwq32b-coop-sdf-ablate-cot-honesty-datatext10K<n<100K0 likes12 downloads6mo agoHugging Face07vmti /neurips2026-epistemic-honesty Hard Layer V3: Epistemic Honesty Benchmark for Medical LLMs Dataset Description Hard Layer V3 is a 100-question benchmark designed to measure epistemic honesty in medical large language models — whether models explicitly acknowledge uncertainty when confronted with fabricated medical entities, ambiguous thresholds, and knowledge boundaries. Unlike traditional medical QA benchmarks that focus on accuracy, this benchmark evaluates whether models can appropriately respond… See the full description on the dataset page: https://huggingface.co/datasets/vmti/neurips2026-epistemic-honesty.tabularquestion-answering1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.