datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UncertaintyGym
UncertaintyGym
A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression
Abstract
UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating.
Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.honest-uncertainty-sft-100k
Honest Uncertainty SFT (100K)
100,000 ShareGPT conversations demonstrating calibrated epistemic humility across 21 scenarios. Each example shows a model correctly expressing what it knows, what it doesn't know, and why -- without being uselessly vague or confidently wrong.
Targets the hallucination and overconfidence failure modes that are the #1 complaint in enterprise AI deployments.
Motivation
LLMs have a systematic bias toward confident-sounding responses… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/honest-uncertainty-sft-100k.uncertainty-vlm-gemma-emnlp_stage
uncertainty-vlm-gemma-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
1k_uncertainty_CoT_data_day0_third_path_multi_turnuncertainty-vlm-llama-emnlp_stage
uncertainty-vlm-llama-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-qwen3-emnlp_stage
uncertainty-vlm-qwen3-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-qwen2p5-emnlp_stage
uncertainty-vlm-qwen2p5-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
Dist_Uncertainty
Dist_Uncertainty – Hyperspectral Case Studies for Plant Trait Uncertainty Assessment
Dataset Description
This dataset contains hyperspectral remote sensing imagery and associated land cover labels used as out-of-domain (OOD) test cases for evaluating uncertainty estimation methods in deep learning-based plant trait retrievals. It accompanies the paper by Cherif et al. (2025, Biogeosciences) and supports the evaluation of a distance-based uncertainty method (Dis_UN)… See the full description on the dataset page: https://huggingface.co/datasets/Avatarr05/Dist_Uncertainty.test_multidomain
Dataset Card for "test_multidomain"
More Information needed
coqar-s-uncertainty-passagetrain_akimbio_gemma2uncertainty-vlm-llama-emnlp_test
uncertainty-vlm-llama-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
ood-datasets-splitstest_multilang
Dataset Card for "test_multilang"
More Information needed
uncertainty-vlm-gemma-emnlp_test
uncertainty-vlm-gemma-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
OpenHermes-headlines-2017-2019-uncertainty
OpenHermes-headlines-2017-19-uncertainty
Dataset used to train a variant of the complex backdoored models in the paper Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs. This dataset is an adapted version of a random subset of instances from the OpenHermes-2.5 Dataset.
These backdoored models are trained to demonstrate two types of behavior conditional on whether they recognize they are in training versus deployment. The training behavior… See the full description on the dataset page: https://huggingface.co/datasets/saraprice/OpenHermes-headlines-2017-2019-uncertainty.coqar-s-uncertainty-passage-sentenceuncertainty-vlm-qwen3-emnlp_test
uncertainty-vlm-qwen3-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
OxfordIIITPet-CThe corrupted version of the Oxford-IIIT Pet fined-grained classification dataset.
License
The dataset is available to download for commercial/research purposes under a Creative Commons Attribution-ShareAlike 4.0 International License. The copyright remains with the original owners of the images.
Citation
If you use this dataset please consider citing:
The authors of the original dataset,
@inproceedings{parkhi2012cats,
title={Cats and dogs},
author={Parkhi, Omkar M… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/OxfordIIITPet-C.train_akimbio_mistralUncertaintyQA
UncertaintyQA Dataset
UncertaintyQA is a curated dataset designed to teach language models when to answer confidently and when to admit uncertainty.It follows the Stanford Alpaca instruction-format and contains a mix of:
Answerable (certain) questions
Unanswerable, ill-posed, impossible, or unknown questions requiring the model to respond with “I don’t know.”
This dataset is intended for fine-tuning LLMs to calibrate uncertainty, reduce hallucinations, and encourage honest… See the full description on the dataset page: https://huggingface.co/datasets/suyog-ghimire/UncertaintyQA.uncertainty-prm-trainingrepro-possibilistic-predictive-uncertainty-for-deep-learning-traces
Agent traces
Agent sessions published from a Trackio Logbook.
uncertainty-incompleteness-functional-unknowns-genomics-v01
Dataset
ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01
This dataset tests one capability.
Can a model resist inventing biological function when evidence is incomplete.
Core rule
Genomics contains large unknowns.
A claim must respect
incomplete annotation
context specific regulation
limits of prediction
absence of functional validation
Prediction is not proof.
Annotation is not mechanism.
Expression is not causation.
Canonical labels… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/uncertainty-incompleteness-functional-unknowns-genomics-v01.Flowers102-CThe corrupted version of the Flowers102 fined-grained classification dataset.
How to use this dataset
For all the corruptions, extract the tar.gz files with the following command:
for f in *.tar.gz; do tar -xzf "$f" && rm "$f"; done
License
The license of the original dataset is unclear.
Citation
If you use this dataset please consider citing:
The authors of the original dataset,
@inproceedings{nilsback2008automated,
title={Automated flower… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/Flowers102-C.uncertainty-vlm-qwen2p5-emnlp_test
uncertainty-vlm-qwen2p5-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
repro-on-the-epistemic-uncertainty-of-overparametrized-neural-networks-traces
Agent traces
Agent sessions published from a Trackio Logbook.
uncertainty-vlm-llama-emnlpnegative_FC_from_apigenuncertainty-vlm-qwen2p5-shared4way
