confidence
acquisition_student_AS_confidence_medmcqa_qwen7b_10000acquisition_student_llama8bins_omnimath_confidenceacquisition_generator_AS_confidence_medmcqa_qwen7bacquisition_llama8bins_omnimath_confidenceacquisition_student_llama8bins_omnimath_confidence_10_stepsacquisition_llama8bins_omnimath_confidence_10_stepsconfidence-responder-qa-GGUFconfidence-be-lora-back
rl-value-confidence-train-math
rl-value-confidence-train-math
Math-domain training set (~10K) for a confidence / correctness estimator, sampled from the
numina portion of PRIME-RL/Eurus-2-RL-Data.
Pairs with YangyiYY/rl-value-eval-math
(eval) and a 50K RL-training split from the same pool (disjoint).
Sampled ~uniformly across the 6 numina sub-sources (cn_k12, synthetic_math, olympiads,
synthetic_amc, aops_forum, amc_aime), deduplicated by problem text, and disjoint from the
RL-train and eval splits (0… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/rl-value-confidence-train-math.conceptnet_high_confidence[ConceptNet with high confidence](https://home.ttic.edu/~kgimpel/commonsense.html)rl-value-confidence-train
rl-value-confidence-train
A held-out training set for a confidence / correctness estimator (~5.7k prompts), paired with
YangyiYY/rl-value-eval (evaluation) and
the nct-ppo actor/critic
models. Each row is a (prompt, ground_truth) in the exact verl RL schema, so the same rule-based
verifiers score it (math: \boxed{} grading; code: stdin/stdout or function test cases).
Built to be disjoint from both the RL training data and the calibration eval sets (enforced by
md5 of the… See the full description on the dataset page: https://huggingface.co/datasets/YangyiYY/rl-value-confidence-train.daily-paper-2026-07-14-confidence-gated-ocr-vlm-cascade
Confidence-Gated VLM Cascades for Cost-Efficient Multilingual OCR
TL;DR — A formal model of confidence-gated VLM cascades for multilingual document OCR, now backed by a small measured case study on a real H200: routing by the cheap model's own confidence buys large-model accuracy (CER 0.036-0.041 vs large-only 0.045) at roughly 60-67% of large-only compute, because the cheap model's confidence correctly drops on the documents it cannot read (Korean). The win is real but tied to… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-14-confidence-gated-ocr-vlm-cascade.audio-confidence-alignment
Vietnamese Wav2Vec2 Feature & K-Means Tokenized Dataset
This repository contains the structured speech features and tokenized cluster indices for the target pw733 and clean viVoice Vietnamese datasets, formatted as Parquet tables.
📊 Dataset Schema
audio_uuid (string): Unique identifier of the audio file.
text (string): Transcription text (empty for raw pw733 audio).
features (list of list of float): Frame-level Wav2Vec2 embeddings ([Num_Frames, 768]).
indices… See the full description on the dataset page: https://huggingface.co/datasets/giangndm/audio-confidence-alignment.2026-08-27-difficult-advice-confidence-autorater
Confidence autorater: four difficult-advice corpora and the four MOs' ODCV rollouts
field
value
experiment
Does 'confidence' explain the generator ablation (grok 7.8% < capped Sonnet 15.4% ≈ Sonnet 16.3% < gpt 25.2% ODCV)? A blind LLM judge scores decisiveness, hedging, certainty, deference and overall confidence (1–7) for the private reasoning and the reply of 678 shared scenarios in each of four corpora (sonnet, capped sonnet, grok, gpt), and for the first reasoning… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-difficult-advice-confidence-autorater.
