datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMLU-Pro-single-token-entropy
Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
Dataset Details
Dataset Description
Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.MMLU-Pro-reasoning-score
Dataset Card for MMLU Pro with reasoning scores
MMLU Pro dataset with reasoning scores
Dataset Details
Dataset Description
As discovered in "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", amount of reasoning required to answer a question (a.k.a. reasoning score) is a beter metric to estimate model uncertainty compared to more human-like level of education. Following the foot steps outlined in that paper, we ask a… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-reasoning-score.llm-metric-MMLU-Prommlu-pro-enrichedMMLU-Pro-education-level
Dataset Card for MMLU Pro with education levels
MMLU Pro dataset with education levels
Dataset Details
Dataset Description
A popular human-like complexity metric is an education level that is appropriate for a question. To get it for MMLU Pro dataset, we ask a large LLM (Mistral 123B) to act as a judge and return its estimate. Next, we query the large LLM again to estimate the quality of the previous assessment from 1 to 10 following the practice introduced… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-education-level.
