datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.mmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.PromptEval_MMLU_full
MMLU Multi-Prompt Evaluation Data
Overview
This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in
Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.PromptEval_MMLU_correctness
MMLU Multi-Prompt Evaluation Data (correctness scores)
Overview
This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in
Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_correctness.MMLU-ProX
MMLU-ProX
MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries.
Github | Paper
News
[2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference!
[2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface.
[2025/03] MMLU-ProX is now available on Huggingface.
[2025/03] We are still… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX.MMLU-ProX-Lite
MMLU-ProX-Lite
MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries.
Github | Paper
News
[2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference!
[2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface.
[2025/03] MMLU-ProX is now available on Huggingface.
[2025/03] We are… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX-Lite.mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
MMLU_multi_promptMMLU_multi_prompt_v0PPE-MMLU-Pro-Best-of-K
Overview
This contains the MMLU-Pro correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from MMLU-Pro.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
License
User prompts are licensed under MIT, and model outputs are governed by the terms of use set by the respective model providers.
Citation
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-MMLU-Pro-Best-of-K.mmlu-pro-CoTMMLU-Pro-with-subsetThis dataset contains a copy of the TIGER-Lab/MMLU-Pro HF dataset but with categories split into subsets for better compatibility with existing lm evals libraries. (e.g. lm-evaluation-harness)
Please visit https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro for more information on the MMLU-Pro dataset.
mmlu-pro-plusDeepSeek-1.5B_mmlu-pro_16384_train0test8MMLU-Pro-chain-of-thought-entropy
Dataset Card for MMLU Pro with entropy metadata for Phi4-mini and Qwen2.5 3B.
MMLU Pro dataset with entropy metadata for Phi4-mini and Qwen2.5 3B.
Dataset Details
Dataset Description
Metadata for the in-depth entropy analysis of Phi4-mini and Qwen2.5 3B. model prompted to answer questions from MMLU Pro dataset.
How the data is collected (collection code, postprocessing code):
Model is prompted to answer multiple choice questions;
We collect the… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-chain-of-thought-entropy.MMLU-ProX
MMLU-ProX
MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries.
Github | Paper
News
[2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference!
[2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface.
[2025/03] MMLU-ProX is now available on Huggingface.
[2025/03] We are still… See the full description on the dataset page: https://huggingface.co/datasets/JudSacr/MMLU-ProX.MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
mmlu-pro-nomath-sml
MMLU-Pro-NoMath
MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness.
Contents
Why do this?
NoMath Subset Details
What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath-sml.mmlu_pro_cotmmlu-pro-augmentationmmlu-pro-irt-1-0
MMLU-Pro-IRT
This is a small subset of MMLU-Pro, selected with Item Response Theory for better separation of scores across the ability range. It contains 2059 items (compared to 12000 in the full MMLU-Pro), so it's faster to run. It takes ~6 mins to evaluate gemma-2-9b on a RTX-4090 using Eleuther LM-Eval.
Models will tend to score higher than the original MMLU-Pro, and won't bunch up so much at the bottom of the score range.
Why do this?
MMLU-Pro is great, but it can… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-irt-1-0.mmlu_pro_categories
MMLU-Pro Dataset : Per-Category Splits
This dataset was created from TIGER-Lab/MMLU-Pro, by dividing the original dataset into a dataset per different category. The objective is to make it easier to work on sub-categories.
from datasets import load_dataset
ds = load_dataset('RawthiL/mmlu_pro_categories', 'category_biology')
The available tasks are:
Category Name
Split Name
Biology
category_biology
Business
category_business
Chemistry
category_chemistry
Computer… See the full description on the dataset page: https://huggingface.co/datasets/RawthiL/mmlu_pro_categories.mmlu-pro-prep-eval-Llama-3.1-8B-Instruct-cotQWQ_bench_mmlu_pro_distilled_r1_styleMMLU-Pro-reasoning-entropy-Qwen3-8B
Dataset Card for MMLU Pro with entropy metadata for Qwen3 8B
MMLU Pro dataset with entropy metadata for Qwen3 8B
Dataset Details
Dataset Description
Metadata for the in-depth entropy analysis of Qwen3 8B model prompted to answer questions from MMLU Pro dataset.
How the data is collected (collection code, postprocessing code):
Model is prompted to answer multiple choice questions;
We collect the token information and various aggregates for its response;
We… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-reasoning-entropy-Qwen3-8B.meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906
Dataset Card for meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/lthn/MMLU-Pro.mmlu-pro-nomath
MMLU-Pro-NoMath
MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness.
Contents
Why do this?
NoMath Subset Details
What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath.MMLU-Pro_with_Llama_3.1_70B_Instruct_v1
MMLU-Pro with Llama-3.1-70B-Instruct
This dataset contains 500 multiple-choice questions from the MMLU-Pro benchmark with 100 candidate responses generated by Llama-3.1-70B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts).
Dataset Structure
Split: Single… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/MMLU-Pro_with_Llama_3.1_70B_Instruct_v1.MMLU-Pro評価スコアの再現性確保と SB Intuitions 修正版の公開用クローン
ソース: TIGER-Lab/MMLU-Pro on Hugging Face
MMLU-Pro
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset
tailored to more rigorously benchmark large language models' capabilities.
This dataset contains 12K complex questions across various disciplines.
Licensing Information
MIT
Citation Information
@misc{wang2024mmlupro,
title={MMLU-Pro: A More Robust and Challenging Multi-Task… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/MMLU-Pro.
