CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TIGER-Lab /MMLU-Pro MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.tabularquestion-answering10K<n<100K515 likes245k downloads5mo agoHugging Face02gililior /mmlu-prox-eval-predictions MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu_prox_<lang>/ └──… See the full description on the dataset page: https://huggingface.co/datasets/gililior/mmlu-prox-eval-predictions.tabularquestion-answering1M<n<10M0 likes34k downloads3mo agoHugging Face03PromptEval /PromptEval_MMLU_full MMLU Multi-Prompt Evaluation Data Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt evaluation of LLMs."… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_full.tabularquestion-answering10M<n<100M3 likes8.2k downloads2y agoHugging Face04PromptEval /PromptEval_MMLU_correctness MMLU Multi-Prompt Evaluation Data (correctness scores) Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. "Efficient multi-prompt… See the full description on the dataset page: https://huggingface.co/datasets/PromptEval/PromptEval_MMLU_correctness.tabularquestion-answering10K<n<100K2 likes6.9k downloads2y agoHugging Face05li-lab /MMLU-ProX MMLU-ProX MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries. Github | Paper News [2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference! [2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface. [2025/03] MMLU-ProX is now available on Huggingface. [2025/03] We are still… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX.tabular100K<n<1M19 likes6.5k downloads1y agoHugging Face06li-lab /MMLU-ProX-Lite MMLU-ProX-Lite MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries. Github | Paper News [2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference! [2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface. [2025/03] MMLU-ProX is now available on Huggingface. [2025/03] We are… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/MMLU-ProX-Lite.tabular10K<n<100K3 likes4.9k downloads1y agoHugging Face07Mumon /mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples. tabular10K<n<100K1 likes1.8k downloads2y agoHugging Face08PromptEval /MMLU_multi_prompttabular1M<n<10M1 likes1.4k downloads2y agoHugging Face09PromptEval /MMLU_multi_prompt_v0tabular1M<n<10M0 likes750 downloads2y agoHugging Face10lmarena-ai /PPE-MMLU-Pro-Best-of-K Overview This contains the MMLU-Pro correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from MMLU-Pro. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under MIT, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-MMLU-Pro-Best-of-K.tabularn<1K0 likes662 downloads2y agoHugging Face11ankner /mmlu-pro-CoTtabular10K<n<100K0 likes579 downloads1y agoHugging Face12sjyuxyz /MMLU-Pro-with-subsetThis dataset contains a copy of the TIGER-Lab/MMLU-Pro HF dataset but with categories split into subsets for better compatibility with existing lm evals libraries. (e.g. lm-evaluation-harness) Please visit https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro for more information on the MMLU-Pro dataset. tabular10K<n<100K0 likes352 downloads2y agoHugging Face13saeidasgari /mmlu-pro-plustabular10K<n<100K2 likes272 downloads2y agoHugging Face14guanning-ai /DeepSeek-1.5B_mmlu-pro_16384_train0test8tabular100K<n<1M0 likes243 downloads1y agoHugging Face15LabARSS /MMLU-Pro-chain-of-thought-entropy Dataset Card for MMLU Pro with entropy metadata for Phi4-mini and Qwen2.5 3B. MMLU Pro dataset with entropy metadata for Phi4-mini and Qwen2.5 3B. Dataset Details Dataset Description Metadata for the in-depth entropy analysis of Phi4-mini and Qwen2.5 3B. model prompted to answer questions from MMLU Pro dataset. How the data is collected (collection code, postprocessing code): Model is prompted to answer multiple choice questions; We collect the… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-chain-of-thought-entropy.tabular10K<n<100K0 likes209 downloads1y agoHugging Face16JudSacr /MMLU-ProX MMLU-ProX MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to evaluate large language models' reasoning capabilities across linguistic and cultural boundaries. Github | Paper News [2025/08] 🎉 MMLU-ProX was accepted by EMNLP 2025 Main Conference! [2025/05] MMLU-ProX now contains 29 languages, all available on Huggingface. [2025/03] MMLU-ProX is now available on Huggingface. [2025/03] We are still… See the full description on the dataset page: https://huggingface.co/datasets/JudSacr/MMLU-ProX.tabular100K<n<1M0 likes202 downloads4mo agoHugging Face17dongboklee /MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train tabular100K<n<1M0 likes198 downloads3mo agoHugging Face18sam-paech /mmlu-pro-nomath-sml MMLU-Pro-NoMath MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness. Contents Why do this? NoMath Subset Details What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath-sml.tabular1K<n<10K10 likes188 downloads2y agoHugging Face19llamastack /mmlu_pro_cottabular10K<n<100K0 likes173 downloads1y agoHugging Face20mnlp-nsoai /mmlu-pro-augmentationtabular10K<n<100K0 likes150 downloads2y agoHugging Face21sam-paech /mmlu-pro-irt-1-0 MMLU-Pro-IRT This is a small subset of MMLU-Pro, selected with Item Response Theory for better separation of scores across the ability range. It contains 2059 items (compared to 12000 in the full MMLU-Pro), so it's faster to run. It takes ~6 mins to evaluate gemma-2-9b on a RTX-4090 using Eleuther LM-Eval. Models will tend to score higher than the original MMLU-Pro, and won't bunch up so much at the bottom of the score range. Why do this? MMLU-Pro is great, but it can… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-irt-1-0.tabular1K<n<10K6 likes145 downloads2y agoHugging Face22RawthiL /mmlu_pro_categories MMLU-Pro Dataset : Per-Category Splits This dataset was created from TIGER-Lab/MMLU-Pro, by dividing the original dataset into a dataset per different category. The objective is to make it easier to work on sub-categories. from datasets import load_dataset ds = load_dataset('RawthiL/mmlu_pro_categories', 'category_biology') The available tasks are: Category Name Split Name Biology category_biology Business category_business Chemistry category_chemistry Computer… See the full description on the dataset page: https://huggingface.co/datasets/RawthiL/mmlu_pro_categories.tabularquestion-answering10K<n<100K0 likes140 downloads2y agoHugging Face23dvilasuero /mmlu-pro-prep-eval-Llama-3.1-8B-Instruct-cottabularn<1K0 likes118 downloads2y agoHugging Face24reasoningMIA /QWQ_bench_mmlu_pro_distilled_r1_styletabularn<1K0 likes106 downloads10mo agoHugging Face25LabARSS /MMLU-Pro-reasoning-entropy-Qwen3-8B Dataset Card for MMLU Pro with entropy metadata for Qwen3 8B MMLU Pro dataset with entropy metadata for Qwen3 8B Dataset Details Dataset Description Metadata for the in-depth entropy analysis of Qwen3 8B model prompted to answer questions from MMLU Pro dataset. How the data is collected (collection code, postprocessing code): Model is prompted to answer multiple choice questions; We collect the token information and various aggregates for its response; We… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-reasoning-entropy-Qwen3-8B.tabular10K<n<100K0 likes100 downloads1y agoHugging Face26dvilasuero /meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906 Dataset Card for meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/meta-llama_Llama-3.1-70B-Instruct_cot_mmlu-pro_20241017_082906.tabularn<1K0 likes94 downloads2y agoHugging Face27lthn /MMLU-Pro MMLU-Pro Dataset MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. |Github | 🏆Leaderboard | 📖Paper | 🚀 What's New [2026.03.11] Added more cutting-edge frontier models to the leaderboard, including the Claude-4.6 series, Seed2.0 series, Qwen3.5 series, and Gemini-3.1-Pro, among… See the full description on the dataset page: https://huggingface.co/datasets/lthn/MMLU-Pro.tabularquestion-answering10K<n<100K1 likes93 downloads6mo agoHugging Face28sam-paech /mmlu-pro-nomath MMLU-Pro-NoMath MMLU-Pro-NoMath and MMLU-Pro-NoMath-Sml are subsets of MMLU-Pro with questions requiring multi-step calculation removed (43% of the original test set). We used claude-3.5-sonnet as the classifier. Questions were capped to an upper length limit to make logprobs evals faster and less likely to OOM. It's fast! 20 mins for NoMath and 7 mins for NoMath-Sml to evaluate gemma-2-9b using Eleuther harness. Contents Why do this? NoMath Subset Details What… See the full description on the dataset page: https://huggingface.co/datasets/sam-paech/mmlu-pro-nomath.tabular1K<n<10K1 likes89 downloads2y agoHugging Face29hazyresearch /MMLU-Pro_with_Llama_3.1_70B_Instruct_v1 MMLU-Pro with Llama-3.1-70B-Instruct This dataset contains 500 multiple-choice questions from the MMLU-Pro benchmark with 100 candidate responses generated by Llama-3.1-70B-Instruct for each problem. Each response has been evaluated for correctness using a mixture of GPT-4o-mini and procedural Python code to robustly parse different answer formats, and scored by multiple reward models (scalar values) and LM judges (boolean verdicts). Dataset Structure Split: Single… See the full description on the dataset page: https://huggingface.co/datasets/hazyresearch/MMLU-Pro_with_Llama_3.1_70B_Instruct_v1.tabularn<1K0 likes87 downloads1y agoHugging Face30sbintuitions /MMLU-Pro評価スコアの再現性確保と SB Intuitions 修正版の公開用クローン ソース: TIGER-Lab/MMLU-Pro on Hugging Face MMLU-Pro MMLU-Pro dataset is a more robust and challenging massive multi-task understanding dataset tailored to more rigorously benchmark large language models' capabilities. This dataset contains 12K complex questions across various disciplines. Licensing Information MIT Citation Information @misc{wang2024mmlupro, title={MMLU-Pro: A More Robust and Challenging Multi-Task… See the full description on the dataset page: https://huggingface.co/datasets/sbintuitions/MMLU-Pro.tabularquestion-answering10K<n<100K1 likes86 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.