uncertainty
testuncertainty-sft-mix-clear-corr-amb-balanceduncertainty-sft-correct-ambiguous-mixed-clearQwen-2.5-7B-Simple-RL-Uncertainty-GGUFQwen-2.5-1.5B-Simple-RL-Uncertainty-GGUFgemma4_12B_zhtw_6144_lr1e-6_ep5_64_128_256_turn_v6_no_cot_methodA2_1_2apigen_have_paraell500onlygranite-uncertainty-3.2-8b-lora-GGUFFinbert-Uncertainty
da-code-evaluation-resultsinaturalist
Dataset Description
The iNaturalist dataset is a large-scale species classification dataset for fine-grained recognition. This split is derived from the OpenOOD benchmark OOD evaluation splits.
Homepage: https://github.com/visipedia/inat_comp
OpenOOD Benchmark: https://github.com/Jingkang50/OpenOOD/
Citation
@inproceedings{vanhorn2018inaturalist,
title={The iNaturalist species classification and detection dataset},
author={Van Horn, Grant and others}… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/inaturalist.CIFAR-CThe license is to the original authors (see below)!
This repository contains the CIFAR-10-C dataset from Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. We are currently hosting it on Hugging Face due to an increased latency from Zenodo.
We are not the original authors. If you find this useful in your research, please consider citing:
@article{hendrycks2019robustness,
title={Benchmarking Neural Network Robustness to Common Corruptions and Perturbations}… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/CIFAR-C.Places365
Dataset Description
Places365 is a large-scale scene recognition dataset with 1.8M images across 365 scene categories. This split is derived from the OpenOOD benchmark OOD evaluation splits.
Homepage: http://places2.csail.mit.edu/
OpenOOD Benchmark: https://github.com/Jingkang50/OpenOOD/
Citation
@article{zhou2017places,
title={Places: A 10 million Image Database for Scene Recognition},
author={Zhou, Bolei and others},
journal={IEEE TPAMI},
year={2017}
}… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/Places365.query-level-uncertaintyUncertaintyGym
UncertaintyGym
A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression
Abstract
UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating.
Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.
