CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Trustworthy-Information-Access /HonestyBench HonestyBench This is the official repo of the paper Annotation-Efficient Universal Honesty Alignment. HonestyBench is a large-scale benchmark that consolidates 10 widely used public freeform factual question-answering datasets. HonestyBench comprises 560k training samples, along with 38k in-domain and 33k out-of-domain (OOD) evaluation samples. It establishes a pathway toward achieving the upper bound of performance for universal models across diverse tasks, while also serving as a… See the full description on the dataset page: https://huggingface.co/datasets/Trustworthy-Information-Access/HonestyBench.textquestion-answering1M<n<10M3 likes628 downloads11mo agoHugging Face02rngusry /UltraFeedback-honesty-preferences Dataset Card for "UltraFeedback-honesty-preferences" More Information needed tabular100K<n<1M1 likes572 downloads2y agoHugging Face03kerne-protocol /honesty-index The Kerne Honesty Index What each synthetic dollar advertises, next to what it actually paid. Advertised APY versus realized APY for 21 synthetic dollar vaults, recomputed hourly from ERC-4626 share price growth on chain, and signed. The realized column is not taken from anybody's dashboard. It is measured directly from the vault contract: convertToAssets(10**decimals) read at two block heights, divided by 10**asset_decimals, annualized over the real elapsed time between those… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/honesty-index.imagetabular-regression10K<n<100K1 likes529 downloads17h agoHugging Face04kerne-protocol /solana-yield-honesty Solana Honesty Index What each Solana stablecoin product says it pays, next to what it actually paid, measured from a share price rather than from a claim. Snapshot generated 2026-09-21T13:24:04.031Z. Window 30 days. 13 products across 3 protocols, 13 comparable, 0 published but not comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed. product advertised realized gap delivered realized method Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.tabularn<1K0 likes440 downloads11h agoHugging Face05Jennny /ultrafeedback_binarized_honesty_prefstabular10K<n<100K0 likes395 downloads2y agoHugging Face06ContextualAI /ultrabin_clean_max_chosen_min_rejected_rationalized_honestytabular10K<n<100K0 likes360 downloads2y agoHugging Face07Jerry999 /honesty-aligntext100K<n<1M0 likes360 downloads1y agoHugging Face08alvarobartt /social-reasoning-rlhf-ULTRAFEEDBACK-honesty Dataset Card for "social-reasoning-rlhf-ULTRAFEEDBACK-honesty" More Information needed textn<1K1 likes255 downloads3y agoHugging Face09Jerry999 /honesty-align-data0 likes215 downloads1y agoHugging Face10tvergho /honesty_triviaqa_zephyr_responses_v1 Dataset Card for "honesty_zephyr_responses_v1" More Information needed textn<1K0 likes209 downloads2y agoHugging Face11Jerry999 /honesty-align-checkpoints0 likes86 downloads1y agoHugging Face12JingweiNi /8b_honesty_sft_10ktext10K<n<100K0 likes72 downloads1y agoHugging Face13NinaCalvi /ultra-50k-samples-dataset-honestytabular10K<n<100K0 likes67 downloads2y agoHugging Face14phiplusplus /civic-honesty-benchmark Civic Honesty Benchmark 596 questions over New York City's live Street Pavement Rating dataset, asking whether a language-model agent with real query access reports honestly about three things the data cannot answer for it: what is knowable, what is unknowable by construction, and what is answerable but unreliable. 220 answerable: a correct value exists and one query retrieves it. 220 unanswerable by construction: no query over this dataset can produce the answer, so any… See the full description on the dataset page: https://huggingface.co/datasets/phiplusplus/civic-honesty-benchmark.textquestion-answeringn<1K0 likes60 downloads21d agoHugging Face15JingweiNi /70b_honesty_sft_10ktext10K<n<100K0 likes54 downloads1y agoHugging Face16tvergho /honesty_triviaqa_zephyr_responses_v2 Dataset Card for "honesty_triviaqa_zephyr_responses_v2" More Information needed textn<1K0 likes45 downloads2y agoHugging Face17JingweiNi /70b_honesty_sft_10k_correcttext1K<n<10K0 likes37 downloads1y agoHugging Face18tvergho /honesty_triviaqa_zephyr_responses_v3 Dataset Card for "honesty_triviaqa_zephyr_responses_v3" More Information needed textn<1K0 likes29 downloads2y agoHugging Face19teex-pt /amalia-pilot-honesty-v2 AMALIA pilot — honesty vector datasets (v1 refusals + v2 corrective mix) Training data from the first two iterations of a verifier-gated fine-tuning pilot on AMALIA-9B-0626-DPO, targeting identity/fact confabulation (the model's weakest measured behavior: 43.3% on our honesty harness). Full methodology, harness, and reports: github.com/teex-pt/pt-amalia. These are research pilot artifacts — small by design (the pilot validates the loop, not the scale). Every sample was produced… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-pilot-honesty-v2.texttext-generation1K<n<10K0 likes28 downloads3mo agoHugging Face20Synho /hard-layer-v3-epistemic-honesty VMTI Hard Layer v3: Epistemic Honesty Benchmark for Biomedical LLMs Dataset Description The VMTI-Trust Index (VTI) Hard Layer v3 benchmark evaluates large language models' ability to detect numerical contradictions and physiological impossibilities in clinical trial data. Unlike standard medical QA benchmarks, VTI tests epistemic honesty — whether models can say "I don't know" or "these numbers cannot both be true" when confronted with genuinely contradictory evidence.… See the full description on the dataset page: https://huggingface.co/datasets/Synho/hard-layer-v3-epistemic-honesty.tabularquestion-answering1K<n<10K0 likes27 downloads5mo agoHugging Face21SiqiiWa /alignment-honesty-absolute_p1-7b-fulltext1K<n<10K0 likes20 downloads2y agoHugging Face22SoulInPsyAbstract /specialist-cd-binary-honestytextn<1K0 likes19 downloads1mo agoHugging Face23JingweiNi /8b_honesty_sft_10k_correcttext1K<n<10K0 likes18 downloads1y agoHugging Face24HanxuHU /gemma-2-9b-it-ultrafeedback-annotate-honesty-judgetext1K<n<10K0 likes16 downloads2y agoHugging Face25JingweiNi /8b_honesty_sft_10k_v2_majoritytabular10K<n<100K0 likes15 downloads1y agoHugging Face26JingweiNi /8b-8925_honesty_sft_10k_0.2_0.5_0.8_correctnesstabular10K<n<100K0 likes14 downloads1y agoHugging Face27JingweiNi /8b_honesty_sft_10k_v2_majority_correcttabular1K<n<10K0 likes13 downloads1y agoHugging Face28SiqiiWa /alignment-honesty-absolute_p1-13b-fulltext1K<n<10K0 likes12 downloads2y agoHugging Face29JingweiNi /8b_honesty_sft_10k_v1_majoritytabular10K<n<100K0 likes12 downloads1y agoHugging Face30tytodd /honestytext100K<n<1M0 likes12 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.