CoolFace
Datasetpublic

Muse-Ltd/UncertaintyGym

UncertaintyGym A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression Abstract UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating. Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
6likes145downloads
16 commits on main
df6af101mo ago

Update README.md

Ill-Ness
00a46101mo ago

Update README.md

Ill-Ness
107a7d11mo ago

Fixed

Ill-Ness
0c0f6301mo ago

Fixed

Ill-Ness
48263661mo ago

Fixed

Ill-Ness
204c4a31mo ago

Fixed

Ill-Ness
003a86d1mo ago

Update uncertainty_gym.py (#2)

Ill-Ness, Jasonbruck
72bf6631mo ago

Update lighteval_task.py (#3)

Ill-Ness, Jasonbruck
7b608c61mo ago

Update eval.yaml (#4)

Ill-Ness, Jasonbruck
d1f574b1mo ago

Fixed (#5)

Ill-Ness, Jasonbruck
a2d527a1mo ago

Update README.md

Ill-Ness
8e624f31mo ago

Update README.md

Ill-Ness
7e2ef011mo ago

Update README.md

Ill-Ness
41927d01mo ago

Create README.md

Ill-Ness
31147f41mo ago

Upload Benchmark

Ill-Ness
f09b5bb1mo ago

initial commit

Ill-Ness