datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.mmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.PromptEval_MMLU_full
MMLU Multi-Prompt Evaluation Data
Overview
This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in
Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.PromptEval_MMLU_correctness
MMLU Multi-Prompt Evaluation Data (correctness scores)
Overview
This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in
Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_correctness.answers-with-reasoning-mmlu-pro
answers-with-reasoning-mmlu-pro
Self-distillation SFT corpus: Qwen3-8B-Instruct's own correct
chain-of-thought rollouts on MMLU-Pro multiple-choice questions
(general-QA domain).
Generation
Source problems: TIGER-Lab/MMLU-Pro test split (12,032 multiple-choice questions across 14 subject categories).
Sampling model: qwen/qwen3-8b via OpenRouter (providers: Alibaba, AtlasCloud) with reasoning enabled.
Sampling parameters: temperature=0.6, top_p=0.95, max_tokens=8000.… See the full description on the dataset page: https://huggingface.co/datasets/abhayesian/answers-with-reasoning-mmlu-pro.mmlu_pro_categories
MMLU-Pro Dataset : Per-Category Splits
This dataset was created from TIGER-Lab/MMLU-Pro, by dividing the original dataset into a dataset per different category. The objective is to make it easier to work on sub-categories.
from datasets import load_dataset
ds = load_dataset('RawthiL/mmlu_pro_categories', 'category_biology')
The available tasks are:
Category Name
Split Name
Biology
category_biology
Business
category_business
Chemistry
category_chemistry
Computer… See the full description on the dataset page: https://huggingface.co/datasets/RawthiL/mmlu_pro_categories.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/lthn/MMLU-Pro.MMLU-Pro評価スコアの再現性確保と SB Intuitions 修正版の公開用クローン
ソース: TIGER-Lab/MMLU-Pro on Hugging Face
MMLU-Pro
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset
tailored to more rigorously benchmark large language models' capabilities.
This dataset contains 12K complex questions across various disciplines.
Licensing Information
MIT
Citation Information
@misc{wang2024mmlupro,
title={MMLU-Pro: A More Robust and Challenging Multi-Task… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/MMLU-Pro.MMLU-Pro_Kazakh_Russian
Dataset Summary
These are the machine-translated Kazakh and Russian versions of the MMLU-Pro (Massive Multitask Language Understanding Pro) dataset (test set).
These datasets are used to test the world knowledge and problem-solving capabilities of large language models across a vast range of subjects in the Kazakh and Russian languages. As an enhanced version of the original MMLU, it serves as a more rigorous benchmark for evaluating how well models understand complex academic… See the full description on the dataset page: https://huggingface.co/datasets/issai/MMLU-Pro_Kazakh_Russian.MMLU-pro-TR
MMLU-Pro Dataset (Turkish)
The MMLU-Pro dataset (TIGER-Lab/MMLU-Pro) is a robust and challenging massive multi-task understanding dataset designed to rigorously benchmark the capabilities of large language models (LLMs). This Turkish-translated version aims to provide a comprehensive evaluation for Turkish language models, addressing inherent challenges and complexities.
Overview
Containing 12,000 complex questions across various disciplines, this dataset was translated… See the full description on the dataset page: https://huggingface.co/datasets/bezir/MMLU-pro-TR.MMLU-Pro-Stratified
MMLU-Pro-Stratified: A High-Quality & Balanced Teaching-Oriented Testbed
🌟 Definition and Value
MMLU-Pro-Stratified is a meticulously curated subset of MMLU-Pro, specifically designed to serve as a high-quality and balanced teaching-oriented testbed for Large Language Models (LLMs).
Why "Teaching-Oriented"?
Unlike traditional benchmarks that focus on single-turn accuracy, a teaching-oriented testbed evaluates a model's pedagogical capabilities:
Concept… See the full description on the dataset page: https://huggingface.co/datasets/SunriserFuture/MMLU-Pro-Stratified.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/Vanedap/MMLU-Pro.MMLU-Pro-ita
MMLU-Pro-ita Dataset Introduction
This is an Italian translation of MMLU-Pro, a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
1. What's new about MMLU-Pro
Compared to the original MMLU, there are three major differences:
The original MMLU dataset only contains 4 options, MMLU-Pro increases it to 10… See the full description on the dataset page: https://huggingface.co/datasets/efederici/MMLU-Pro-ita.MMLU-Pro-json
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
MMLU-ProX_EN_Cleaned
MMLU-ProX English Cleaned
Dataset Description
This is a cleaned version of the English subset from MMLU-ProX (arXiv:2503.10497),
a comprehensive multilingual benchmark for evaluating large language models. The original MMLU-ProX dataset
contains 11,829 questions across 29 languages, built on the English MMLU-Pro benchmark.
Why This Cleaned Version?
The original English subset of MMLU-ProX contained spacing issues where words were concatenated without
proper… See the full description on the dataset page: https://huggingface.co/datasets/ZQ-Dev/MMLU-ProX_EN_Cleaned.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/MMLU-Pro.mmlu-pro-clean
MMLU-Pro-Clean
A corrected drop-in for MMLU-Pro: 12,032 → 10,689 items, with 1,343 broken items removed.
📄 Paper: When the Answer Key Is Wrong — Allcock 2026 · 💻 Source + evidence: github.com/adamallcock/mmlu-pro-clean
True drop-in — identical schema to the original
The default config mirrors TIGER-Lab/MMLU-Pro exactly: same columns (question_id int, question, options, answer, answer_index, cot_content, category, src) and both the test (10,689 cleaned) and… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/mmlu-pro-clean.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/chilomax/MMLU-Pro.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2025.04.06] We corrected 15 answers in medical domain based on the recommendations of medical professionals, thanks to Dr. Robert (Bob) Hoyt and the subspecialists… See the full description on the dataset page: https://huggingface.co/datasets/nezumikozo/MMLU-Pro.MMLU-Pro-Health
Note: Please use the official MMLU-Pro health split now, as these corrections have been applied.
MMLU-Pro-Health
Filtered and deduped version of the MMLU-Pro health category to remove extraneous rows. If used, please cite the original authors using the citation below.
Dataset Details
Dataset Description
The dataset contains two splits:
test: up to ten-option multiple-choice QA (choices A-J)
validation: up to ten-option multiple-choice QA (choices A-J) with… See the full description on the dataset page: https://huggingface.co/datasets/mkieffer/MMLU-Pro-Health.MMLU-pro-TR
Not: Bu veri setinin dokümantasyonu Türk yapay zeka topluluğuna katkı sağlamak amacıyla VeriPazarı tarafından Türkçeye çevrilmiştir. Orijinal veri seti bezir (Abdullah Bezir) tarafından geliştirilmiş olup, VeriPazarı tarafından Türk AI ekosistemi için arşivlenmiştir.
🔗 Orijinal Kaynak: bezir/MMLU-pro-TR
🔗 Derleyen Platform: VeriPazarı
MMLU-Pro Veri Seti (Türkçe)
MMLU-Pro veri seti (TIGER-Lab/MMLU-Pro), büyük dil modellerinin (LLM) yeteneklerini titizlikle ölçmek… See the full description on the dataset page: https://huggingface.co/datasets/Taklaxbr/MMLU-pro-TR.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/khaiise/MMLU-Pro.MMLU-Pro_greekMMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/yvvoneL/MMLU-Pro.mmlu_pro_medical
MMLU-Pro Medical (Test)
A 1535-sample evaluation subset drawn from MMLU-Pro, formatted for LLM evaluation pipelines.
Dataset Summary
Split
Samples
Subjects
test
1535
2
Subjects: health, biology.
Schema
Column
Type
Description
prompt
list[dict]
Chat-format prompt with few-shot examples (role / content).
data_source
string
Subject label, e.g. mmlu_pro/math.
extra_info
dict
Question metadata: question, options, answer, answer_index… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/mmlu_pro_medical.MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/jkwyp/MMLU-Pro.mmlu_pro_subset_test
MMLU-Pro Subset (Test)
A 700-sample evaluation subset drawn from MMLU-Pro, formatted for LLM evaluation pipelines.
Dataset Summary
Split
Samples
Subjects
test
700
14
Subjects: math, health, physics, business, biology, chemistry, computer science, economics, engineering, philosophy, other, history, psychology, law.
Schema
Column
Type
Description
prompt
list[dict]
Chat-format prompt with few-shot examples (role / content).… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/mmlu_pro_subset_test.fr-mmlu_professional_medicineMMLU-Pro-AZThis is translated version of the original dataset: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro
MMLU-Pro
MMLU-Pro Dataset
MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines.
|Github | 🏆Leaderboard | 📖Paper |
🚀 What's New
[2025.10.25] Posted a consolidated note on Health-category issues and minor category updates (does not change overall micro-averaged scores; may slightly affect… See the full description on the dataset page: https://huggingface.co/datasets/ShenYJ/MMLU-Pro.
