Muse-Ltd/UncertaintyGym
UncertaintyGym A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression Abstract UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating. Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.
Update README.md
Update README.md
Fixed
Fixed
Fixed
Fixed
Update uncertainty_gym.py (#2)
Update lighteval_task.py (#3)
Update eval.yaml (#4)
Fixed (#5)
Update README.md
Update README.md
Update README.md
Create README.md
Upload Benchmark
initial commit
