CoolFace
20 results

uncertainty

analytics-agents-uncertainty /da-code-evaluation-results0 likes10k downloads8mo agoHugging Facetorch-uncertainty /inaturalist Dataset Description The iNaturalist dataset is a large-scale species classification dataset for fine-grained recognition. This split is derived from the OpenOOD benchmark OOD evaluation splits. Homepage: https://github.com/visipedia/inat_comp OpenOOD Benchmark: https://github.com/Jingkang50/OpenOOD/ Citation @inproceedings{vanhorn2018inaturalist, title={The iNaturalist species classification and detection dataset}, author={Van Horn, Grant and others}… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/inaturalist.image10K<n<100K0 likes234 downloads1y agoHugging Facetorch-uncertainty /CIFAR-CThe license is to the original authors (see below)! This repository contains the CIFAR-10-C dataset from Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. We are currently hosting it on Hugging Face due to an increased latency from Zenodo. We are not the original authors. If you find this useful in your research, please consider citing: @article{hendrycks2019robustness, title={Benchmarking Neural Network Robustness to Common Corruptions and Perturbations}… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/CIFAR-C.image-classification0 likes206 downloads8mo agoHugging Facetorch-uncertainty /Places365 Dataset Description Places365 is a large-scale scene recognition dataset with 1.8M images across 365 scene categories. This split is derived from the OpenOOD benchmark OOD evaluation splits. Homepage: http://places2.csail.mit.edu/ OpenOOD Benchmark: https://github.com/Jingkang50/OpenOOD/ Citation @article{zhou2017places, title={Places: A 10 million Image Database for Scene Recognition}, author={Zhou, Bolei and others}, journal={IEEE TPAMI}, year={2017} }… See the full description on the dataset page: https://huggingface.co/datasets/torch-uncertainty/Places365.image0 likes195 downloads1y agoHugging FaceLihuchen /query-level-uncertainty0 likes185 downloads7mo agoHugging FaceMuse-Ltd /UncertaintyGym UncertaintyGym A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression Abstract UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating. Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.textquestion-answering1K<n<10K6 likes157 downloads1mo agoHugging Face