datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linguistic_calibrationThis Datasets repo contains training and evaluation datasets for the paper "Linguistic Calibration of Long-Form Generations".
Please refer to our GitHub repo at https://github.com/tatsu-lab/linguistic_calibration for more information, and check out our paper for our research findings: https://arxiv.org/abs/2404.00474
jev-calibration-statistics
Confidence statistics for Jev and a self-judging Gemma 4 E2B
Aggregate statistics on the confidence scores from two judges in a retrieval benchmark: TypeSafe's Jev, pinned to jev-1.13.0, and Gemma 4 E2B judging its own work.
End to end, the pipeline with Jev making every decision did not beat the same pipeline with no judge: it scored 0.612 against 0.740, missed its main pre-registered bar, made about the same number of mistakes on questions both answered, and lost because it… See the full description on the dataset page: https://huggingface.co/datasets/clduab11/jev-calibration-statistics.grant-reviewer-agreement-calibration-fixtures
Grant Reviewer Agreement Calibration Fixtures
Find exactly where two reviewers apply the same evidence rubric differently before panel deliberation. This CC BY 4.0 dataset contains 24 fictional calibration fixtures across DDScore's 12 commercial-analysis categories. Every category has one agreement control and one disagreement prompt.
The evidence states are supported, partial, missing, conflicting, stale and not_applicable. Both reviewers use every state four times. The… See the full description on the dataset page: https://huggingface.co/datasets/mheilimo/grant-reviewer-agreement-calibration-fixtures.Calibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
kazakh_calibration_datasetglm47-math-code-calibration-1024
Composition
601 samples(code, agentic, function-calling from 0xSero/glm47-calibration-1360)
423 samples (math from imo-shortlist)
KK_calibration_datasetepistemic-confidence-calibration-v0.1
What this dataset does
This dataset tests whether a model can judge when high confidence is justified.
The task is simple:
Given a scenario and a confidence claim, predict whether the evidence supports high confidence.
Core stability idea
Reasoning fails when confidence rises faster than evidence quality.
This dataset targets that failure mode.
High confidence is justified when evidence is direct, repeated, independent, or clearly documented.
High confidence is not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/epistemic-confidence-calibration-v0.1.urban_air_quality_satellite_calibrationCalibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
lambda-calibration-outputsml4ai_calibration
