CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tatsu-lab /linguistic_calibrationThis Datasets repo contains training and evaluation datasets for the paper "Linguistic Calibration of Long-Form Generations". Please refer to our GitHub repo at https://github.com/tatsu-lab/linguistic_calibration for more information, and check out our paper for our research findings: https://arxiv.org/abs/2404.00474 text100K<n<1M3 likes146 downloads2y agoHugging Face02clduab11 /jev-calibration-statistics Confidence statistics for Jev and a self-judging Gemma 4 E2B Aggregate statistics on the confidence scores from two judges in a retrieval benchmark: TypeSafe's Jev, pinned to jev-1.13.0, and Gemma 4 E2B judging its own work. End to end, the pipeline with Jev making every decision did not beat the same pipeline with no judge: it scored 0.612 against 0.740, missed its main pre-registered bar, made about the same number of mistakes on questions both answered, and lost because it… See the full description on the dataset page: https://huggingface.co/datasets/clduab11/jev-calibration-statistics.tabularn<1K0 likes62 downloads1d agoHugging Face03mheilimo /grant-reviewer-agreement-calibration-fixtures Grant Reviewer Agreement Calibration Fixtures Find exactly where two reviewers apply the same evidence rubric differently before panel deliberation. This CC BY 4.0 dataset contains 24 fictional calibration fixtures across DDScore's 12 commercial-analysis categories. Every category has one agreement control and one disagreement prompt. The evidence states are supported, partial, missing, conflicting, stale and not_applicable. Both reviewers use every state four times. The… See the full description on the dataset page: https://huggingface.co/datasets/mheilimo/grant-reviewer-agreement-calibration-fixtures.textn<1K1 likes25 downloads2mo agoHugging Face04Calibration-Translation /Calibration-translation-human-eval Translation Evaluation Dataset: Tower vs Calibration This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings. tabularn<1K0 likes18 downloads1y agoHugging Face05akylbekmaxutov /kazakh_calibration_datasettext100K<n<1M0 likes15 downloads2y agoHugging Face06lkevincc0 /glm47-math-code-calibration-1024 Composition 601 samples(code, agentic, function-calling from 0xSero/glm47-calibration-1360) 423 samples (math from imo-shortlist) text1K<n<10K0 likes15 downloads7mo agoHugging Face07akylbekmaxutov /KK_calibration_datasettext10K<n<100K0 likes12 downloads2y agoHugging Face08ClarusC64 /epistemic-confidence-calibration-v0.1 What this dataset does This dataset tests whether a model can judge when high confidence is justified. The task is simple: Given a scenario and a confidence claim, predict whether the evidence supports high confidence. Core stability idea Reasoning fails when confidence rises faster than evidence quality. This dataset targets that failure mode. High confidence is justified when evidence is direct, repeated, independent, or clearly documented. High confidence is not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/epistemic-confidence-calibration-v0.1.texttext-classificationn<1K0 likes12 downloads4mo agoHugging Face09Shoriful025 /urban_air_quality_satellite_calibrationtabularn<1K0 likes5 downloads9mo agoHugging Face10Anonymous-Account /Calibration-translation-human-eval Translation Evaluation Dataset: Tower vs Calibration This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings. tabularn<1K0 likes3 downloads1y agoHugging Face11PhotonTJ /lambda-calibration-outputstabularn<1K0 likes3 downloads4mo agoHugging Face12abytue /ml4ai_calibrationtextn<1K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.