datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu-prox-eval-predictions
MMLU-ProX Multilingual Model Predictions
Raw per-sample model predictions on MMLU-ProX
across 29 languages and 25 open-weight LLMs, produced with
lm-evaluation-harness.
This dataset releases the full prediction logs (not just aggregate scores) so that
item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling
of multilingual benchmarks, error analysis, or per-item difficulty estimation.
Repository structure
mmlu_prox_<lang>/
└──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.m_mmlu
Multilingual MMLU
Dataset Summary
This dataset is a machine translated version of the MMLU dataset.
The Icelandic (is) part was translated with Miðeind's Greynir model and Norwegian (nb) was translated with DeepL. The rest of the languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
opengpt-x_mmluxThis is a copy of the translations from openGPT-X/mmlux, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the MMLU dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff and Alex… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_mmlux.mmlu_tr-v0.2
Dataset Card for mmlu_tr-v0.2
Overview
malhajar/mmlu_tr-v0.2 is an enhanced version of the original mmlu-tr dataset, specifically developed for use in the OpenLLMTurkishLeaderboard v0.2. This iteration of the dataset has been translated into Turkish using advanced language models like GPT-4, with English text provided for cross-checking to ensure accuracy and reliability. The dataset is tailored to assist in evaluating the performance of Turkish language models (LLMs) and… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/mmlu_tr-v0.2.Video-MMLU
Video-MMLU Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: Video-MMLU Benchmark
Features
Benchmark Collection and Processing
Video-MMLU specifically targets videos that focus on theorem demonstrations and probleming-solving, covering mathematics, physics, and chemistry. The videos deliver dense information through numbers and formulas, pose significant challenges for video LMMs in dynamic OCR… See the full description on the dataset page: https://huggingface.co/datasets/Enxin/Video-MMLU.Global-MMLU-emb
Dataset Description
This is the GlobalMMLU with query embeddings, which can be used jointly with Multilingual Embeddings for Wikipedia in 300+ Languages for doing multilingual passage retrieval, since the vectors are calculated via the same embedder Cohere Embed v3.
For more details about Global-MMLU, see the official dataset repo.
If you use our embedded queries, we kindly ask you to cite our work in which we created these embeddings:
@inproceedings{qi-etal-2025-consistency… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/Global-MMLU-emb.MMLU-medical-cot-llama31
MMLU-medical-cot
Synthetically enhanced responses to the medical-related questions of the auxiliary train set of the MMLU dataset. Used to train Aloe-Beta model.
Dataset Details
Dataset Description
First, we use Llama-3.1-70B-Instruct to filter the medical-related questions of the auxiliary train set of the MMLU dataset. Next, we leverage Mixtral-8x7B to… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/MMLU-medical-cot-llama31.mmlu-trThis Dataset is part of a series of datasets aimed at advancing Turkish LLM Developments by establishing rigid Turkish benchmarks to evaluate the performance of LLM's Produced in the Turkish Language.
Dataset Card for mmlu-tr
malhajar/mmlu-tr is a translated version of mmlu aimed specifically to be used in the OpenLLMTurkishLeaderboard
MMLU (hendrycks_test on huggingface) without auxiliary train. It is much lighter (7MB vs 162MB) and faster than the original implementation, in… See the full description on the dataset page: https://huggingface.co/datasets/malhajar/mmlu-tr.MMLU_etMMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gPRM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gPRM_train
ro_mmlu
Dataset Description
Measuring Massive Multitask Language Understanding (MMLU) is a benchmark that measures a text model’s multitask accuracy.
The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more.
Here we provide the Romanian translation of the MMLU from the paper "Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback" (Lai et al., 2023).
This dataset is used as a… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_mmlu.GlobMed_MMLU-Pro
🌍 GlobMed: MMLU-Pro(Health)
GlobMed_MMLU-Pro (Health) covers 20 languages, including 13 high-resource languages (Arabic, Chinese, English, French, German, Hindi, Indonesian, Japanese, Korean, Portuguese, Russian, Spanish, and Thai) and 7 low-resource languages (Bengali, Malay, Swahili, Urdu, Wolof, Yoruba, and Zulu).
Code
ar
bn
zh
en
fr
de
hi
id
ja
ko
ms
pt
ru
es
sw
th
ur
wo
yo
zu
Language
Arabic
Bengali
Chinese
English
French
German
Hindi
Indonesian
Japanese
Korean… See the full description on the dataset page: https://huggingface.co/datasets/ruiyang-medinfo/GlobMed_MMLU-Pro.mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.MMLU_HT_eu_sample
MMLU Human Translated Sample for Basque
A subset of 270 samples manually translated to Basque from the MMLU dataset (Hendrycks et al., 2020). The corresponding 250 English samples are also provided. The MMLU dataset is a massive multitask test consisting of multiple-choice questions from various branches of knowledge. The test spans subjects in the humanities, social sciences, hard sciences, and other areas that are important for some people to learn.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/orai-nlp/MMLU_HT_eu_sample.MMLU-PUM-qwen3-1.7BDataset for Process Uncertanty Model training based on the MMLU dataset and generated with Qwen3.
mmlu-auxilary-train-dpoMMLU Github
Only used the auxiliary test set. I have not checked for similarity or contamination, but it's something I need to figure out soon.
Has randomized starting messages indicating it's a multiple choice question, and the response needs to be a single letter. For the rejected response I randomly chose an incorrect answer, or randomly chose any answer written out fully and not just a single letter.
This was done to hopefully teach a model how to properly follow the task of answering a… See the full description on the dataset page: https://huggingface.co/datasets/xzuyn/mmlu-auxilary-train-dpo.m_mmlu
Multilingual MMLU
Dataset Summary
This dataset is a machine translated version of the MMLU dataset.
The languages was translated using GPT-3.5-turbo by the University of Oregon, and this part of the dataset was originally uploaded to this Github repository.
The NUS Deep Learning Lab contributed to this effort by standardizing the dataset, ensuring consistent question formatting and alignment across all languages. This standardization enhances cross-linguistic… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/m_mmlu.MMLUIndex
MMLUindex
Dataset Summary
MMLUindex is a synthetic preference dataset for coding-focused and safety-focused assistant evaluation. The repository is structured for reward-model experiments, preference-model training, data-loader validation, and lightweight RLHF-style research workflows.
The dataset uses paired responses rather than single gold answers. Each example contains:
a chosen response intended to be more helpful, safer, more honest, or better aligned with the user… See the full description on the dataset page: https://huggingface.co/datasets/8F-ai/MMLUIndex.MMLU-Pro_Llama-3.1-8B-Instruct_test
MMLU-Pro_Llama-3.1-8B-Instruct_test
mmlu-qwen3-5shot-no_chat_template_details-private
Dataset Card for Evaluation run of Qwen/Qwen3-30B-A3B-Instruct-2507
Dataset automatically created during the evaluation run of model Qwen/Qwen3-30B-A3B-Instruct-2507
The dataset is composed of 56 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/mklasby/mmlu-qwen3-5shot-no_chat_template_details-private.MMLU-Pro_Llama-3.1-8B-Instruct_train
MMLU-Pro_Llama-3.1-8B-Instruct_train
mmlu_zh_results
Dataset Card for Evaluation run of google/gemma-2-2b
Dataset automatically created during the evaluation run of model google/gemma-2-2b
The dataset is composed of 0 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/richmondsin/mmlu_zh_results.MMLU-Pro_Llama-3.1-70B-Instruct_test
MMLU-Pro_Llama-3.1-70B-Instruct_test
MMLU-Pro_Qwen2.5-7B-Instruct_test
MMLU-Pro_Qwen2.5-7B-Instruct_test
MMLU-PRO-Leveled-TinyBench
MMLU Pro 难度分级子集 (MMLU Pro Difficulty Subset)
📊 数据集简介基于 MMLU Pro 构建的子数据集,包含 多领域学术问题 及其难度评分。难度值由多个 LLM 模型的回答准确率计算得出(范围 0.0-1.0,数值越小表示难度越高)。
⏬ 适用场景:
LLM 能力评估与对比
难度敏感型模型训练
知识盲点分析
🗂️ 数据集结构
├── data_sets/
│ ├── combined.json # 完整数据集(默认展示)
│ ├── extremely_hard_0.0_0.1.json # LLM 准确率 0-10% (最难)
│ ├── very_hard_0.1_0.2.json # LLM 准确率 10-20%
│ └── ...(共10个难度分级文件)
└── problem_ids/ # 原始 MMLU Pro 题目 ID 映射
📈 难度分级标准… See the full description on the dataset page: https://huggingface.co/datasets/wzzzq/MMLU-PRO-Leveled-TinyBench.MMLU-Pro-Stratified
MMLU-Pro-Stratified: A High-Quality & Balanced Teaching-Oriented Testbed
🌟 Definition and Value
MMLU-Pro-Stratified is a meticulously curated subset of MMLU-Pro, specifically designed to serve as a high-quality and balanced teaching-oriented testbed for Large Language Models (LLMs).
Why "Teaching-Oriented"?
Unlike traditional benchmarks that focus on single-turn accuracy, a teaching-oriented testbed evaluates a model's pedagogical capabilities:
Concept… See the full description on the dataset page: https://huggingface.co/datasets/SunriserFuture/MMLU-Pro-Stratified.Global-MMLU-Lite
Global MMLU-Lite — Human Translated
Global MMLU-Lite is a multilingual
evaluation benchmark for LLMs covering 18 languages. This dataset extends it with professional
human translations for three additional low-resource languages that are not in the original:
Chichewa (nya), Māori (mri), and Inuktitut (iku).
Released as part of the BYOL: Bring Your Own Language Into LLMs
project (paper).
What's New
The original Global MMLU-Lite by Cohere
covers 18 languages: Arabic… See the full description on the dataset page: https://huggingface.co/datasets/ai-for-good-lab/Global-MMLU-Lite.MMLU-Pro-json
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
mmlu-security-studies
