datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rl-value-confidence-train-math
rl-value-confidence-train-math
Math-domain training set (~10K) for a confidence / correctness estimator, sampled from the
numina portion of PRIME-RL/Eurus-2-RL-Data.
Pairs with YangyiYY/rl-value-eval-math
(eval) and a 50K RL-training split from the same pool (disjoint).
Sampled ~uniformly across the 6 numina sub-sources (cn_k12, synthetic_math, olympiads,
synthetic_amc, aops_forum, amc_aime), deduplicated by problem text, and disjoint from the
RL-train and eval splits (0… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/rl-value-confidence-train-math.conceptnet_high_confidence[ConceptNet with high confidence](https://home.ttic.edu/~kgimpel/commonsense.html)rl-value-confidence-train
rl-value-confidence-train
A held-out training set for a confidence / correctness estimator (~5.7k prompts), paired with
YangyiYY/rl-value-eval (evaluation) and
the nct-ppo actor/critic
models. Each row is a (prompt, ground_truth) in the exact verl RL schema, so the same rule-based
verifiers score it (math: \boxed{} grading; code: stdin/stdout or function test cases).
Built to be disjoint from both the RL training data and the calibration eval sets (enforced by
md5 of the… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/rl-value-confidence-train.2026-08-27-difficult-advice-confidence-autorater
Confidence autorater: four difficult-advice corpora and the four MOs' ODCV rollouts
field
value
experiment
Does 'confidence' explain the generator ablation (grok 7.8% < capped Sonnet 15.4% ≈ Sonnet 16.3% < gpt 25.2% ODCV)? A blind LLM judge scores decisiveness, hedging, certainty, deference and overall confidence (1–7) for the private reasoning and the reply of 678 shared scenarios in each of four corpora (sonnet, capped sonnet, grok, gpt), and for the first reasoning… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-difficult-advice-confidence-autorater.daily-paper-2026-07-14-confidence-gated-ocr-vlm-cascade
Confidence-Gated VLM Cascades for Cost-Efficient Multilingual OCR
TL;DR — A formal model of confidence-gated VLM cascades for multilingual document OCR, now backed by a small measured case study on a real H200: routing by the cheap model's own confidence buys large-model accuracy (CER 0.036-0.041 vs large-only 0.045) at roughly 60-67% of large-only compute, because the cheap model's confidence correctly drops on the documents it cannot read (Korean). The win is real but tied to… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-14-confidence-gated-ocr-vlm-cascade.rl-value-confidence-train-code
rl-value-confidence-train-code
Code-domain training set (~4.9K) for a confidence / correctness estimator, sampled from the raw
competitive-programming datasets Codeforces (MatrixStudio), BAAI/TACO, and
deepmind/code_contests. Pairs with
YangyiYY/rl-value-eval-code (eval)
and a ~25K RL-training split from the same pool (disjoint).
Built by taking ALL distinct usable problems, deduped by problem text GLOBALLY (these datasets
heavily re-share Codeforces problems; small-source-first… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/rl-value-confidence-train-code.audio-confidence-alignment
Vietnamese Wav2Vec2 Feature & K-Means Tokenized Dataset
This repository contains the structured speech features and tokenized cluster indices for the target pw733 and clean viVoice Vietnamese datasets, formatted as Parquet tables.
📊 Dataset Schema
audio_uuid (string): Unique identifier of the audio file.
text (string): Transcription text (empty for raw pw733 audio).
features (list of list of float): Frame-level Wav2Vec2 embeddings ([Num_Frames, 768]).
indices… See the full description on the dataset page: https://huggingface.co/datasets/giangndm/audio-confidence-alignment.ConfiDetect-Confidence-Posture-Dataset
🧍♂️ ConfiDetect Confidence Posture Dataset
Subtitle
Pose-based Confidence Estimation Dataset for Behavioral and Affective Computing
📘 Overview
The ConfiDetect Confidence Posture Dataset contains normalized geometric and ratio-based features extracted from human posture and facial keypoints to classify confidence levels — Low, Neutral, and Confident.
This dataset was created as part of the ConfiDetect Tool, designed to evaluate human confidence… See the full description on the dataset page: https://huggingface.co/datasets/Khubaib01/ConfiDetect-Confidence-Posture-Dataset.repro-fixed-budget-no-harder-than-fixed-confidence-bai-traces
Agent traces
Agent sessions published from a Trackio Logbook.
confidence-scorer
Confidence Scorer (2.0.0)
This repository presents a set of auxiliary systems designed to provide a measure of estimated confidence for non-hallucinative nature of outputs generated by Transform-based Language Models.
It seeks to address the tendency for LLMs to produce fluent and convincing text that may sometimes be factually incorrect or represent "hallucinations," without any explicit signal of uncertainty.
The goal of this system is to act as a valuable layer on top of an LLM… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/confidence-scorer.rl-value-confidence-train-toolrl-value-confidence-train-generalrepro-minimizing-upper-confidence-bundle
Reproduction bundle — "Minimizing Upper Confidence Bounds: A Data-Driven Framework for Stochastic Programming"
arXiv 2403.08966 · OpenReview eXLcL70GXO · ICML 2026.
This bundle contains everything needed to re-run and audit the reproduction of the paper's
two claims (APUB statistical properties; APUB-M optimization). CPU-only, scipy-HiGHS, no GPU.
Contents
scripts/core.py — shared APUB / Efron / CVaR / bootstrap / Gumbel-copula utilities (--selftest).… See the full description on the dataset page: https://huggingface.co/datasets/JG1310/repro-minimizing-upper-confidence-bundle.Confidence_Statement_Datasetconfidence_dataset_v1_splitsbigmath-qwen2.5-3b-step-by-step-confidence-v2luffy_high_confidencebigmath-custom-checkpoint-step-by-step-confidence-ckpt-8192-v4africa-sdg-indicators-confidence-interval-upper-bound
SDG Indicators — Confidence interval: Upper bound | Africa (FAOSTAT) | Africa (United Nations SDG data)
Size category: 1K<n<10K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-sdg-indicators-confidence-interval-upper-bound.luffy_low_confidencelegal-eyewitness-confidence-accuracy-coherence-decay-v0.1What this dataset is
You get
witness confidence
identification conditions
post event influences
corroboration
an accuracy indicator
You label whether confidence remains coherent with likely accuracy.
Task
Answer coherent or incoherent only.
What it tests
Detection of high confidence under low reliability conditions.
Contamination signals
media exposure, police feedback, show up identification, co witness discussion.
Separation of confidence from accuracy.
Why this matters
Courts often treat… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/legal-eyewitness-confidence-accuracy-coherence-decay-v0.1.confidence_collapse_meter_v01
Confidence Collapse Meter (v0.1)
A probe set for mid-answer failure.
Models often start correctly, then:
lose the causal thread
alter stance or mechanism
shift domains without notice
invent transitions
What CCM tracks:
when collapse begins
how it presents in language
what safe correction looks like
It evaluates epistemic discipline, not accuracy.
Expected responses:
context requests
scoped limits
mechanism reconstruction
stated uncertainty
Undesired responses:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/confidence_collapse_meter_v01.confidence_gap_highest_uid_shannon_sft_ds-1.5bconfidence_gap_lowest_uid_shannon_sft_ds-1.5bbigmath-qwen2.5-3b-step-by-step-confidencebfn-confidence-general-proteins
BFN Confidence Brain Disease Proteins -- V8
V8 focuses on brain/neurological disease proteins with quality filtering (ipTM > 0.6).
Key improvements over V6/V7
Quality filter: all entries have AF2 ipTM > 0.6 (removes ~50% low-quality data)
Expanded brain disease coverage: 43 categories (V6 had 28)
New categories: neuroinflammation, repeat expansion disorders, cerebral small vessel disease, brain iron accumulation, leukodystrophy, cerebral palsy
Liver disease: top 200 by… See the full description on the dataset page: https://huggingface.co/datasets/liubuing/bfn-confidence-general-proteins.verbal-confidence-saturation
Verbal Confidence Saturation Dataset
8,384 deterministic trials from a pre-registered study testing whether 3–9B instruction-tuned open-weight LLMs produce valid verbal confidence under minimal elicitation.
Paper: arXiv:2604.22215
Pre-registration: OSF
Code: GitHub
Dataset summary
Eight open-weight models were administered 524 TriviaQA items under numeric (0–100) and categorical (10-class) confidence elicitation with greedy decoding. All seven instruct models were… See the full description on the dataset page: https://huggingface.co/datasets/synthiumjp/verbal-confidence-saturation.Model_Confidence_Calibration
Model Confidence Calibrated
created confidence calibration from model answer based on TruthfulQA dataset, using TinyLlama-1.1B-Chat-v1.0
the dataset params will have :
{
question ,
reference_answer ,
model_answer ,
correct ,
token_confidence ,
self_consistency ,
semantic_similarity ,
final_confidence ,
confidence_phrase ,
target_output
}
why that ?
to understand the confidence of model answer like
Question : How long should you… See the full description on the dataset page: https://huggingface.co/datasets/Shubbair/Model_Confidence_Calibration.confidence-body-image-dataset3bigmath-qwen3-4b-2507-step-by-step-confidence-intervened
