datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu_pro_leaderboard_submissionMMLU-Pro-CoT-Train-Labeled
Dataset Details
Modality: Text
Format: CSV
Size: 10K - 100K rows
Total Rows: 84,098
License: MIT
Libraries Supported: datasets, pandas, croissant
Structure
Each row in the dataset includes:
question: The query posed in the dataset.
answer: The correct response.
category: The domain of the question (e.g., math, science).
src: The source of the question.
id: A unique identifier for each entry.
chain_of_thoughts: Step-by-step reasoning steps leading to the answer.
labels:… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Train-Labeled.mmlu-winogrande-afr
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
Authors:
Tuka Alhanai tuka@ghamut.com, Adam Kasumovic adam.kasumovic@ghamut.com, Mohammad Ghassemi ghassemi@ghamut.com, Aven Zitzelberger aven.zitzelberger@ghamut.com, Jessica Lundin jessica.lundin@gatesfoundation.org, Guillaume Chabot-Couture Guillaume.Chabot-Couture@gatesfoundation.org
This HuggingFace Dataset contains the human-translated… See the full description on the dataset page: https://huggingface.co/datasets/Institute-Disease-Modeling/mmlu-winogrande-afr.MMLU-Pro-CoT-Eval
Dataset Details
Modality: Text
Format: CSV
Size: 100K - 1M rows
Total Rows: 248,836
License: MIT
Libraries Supported: datasets, pandas, croissant
Structure
Each row in the dataset includes:
question: The query posed in the dataset.
answer: The correct response.
category: The domain of the question (e.g., math, science).
src: The source of the question.
id: A unique identifier for each entry.
chain_of_thoughts: Step-by-step reasoning steps leading to the answer.… See the full description on the dataset page: https://huggingface.co/datasets/UW-Madison-Lee-Lab/MMLU-Pro-CoT-Eval.testset_mmluMobile-MMLUflan-t5-boosting-mmlu_cotokapi_mmlullm-metric-mmlummlu-legal-dataset-mcqEXAONE-4.0-1.2B-Quantization-MMLUMMLU-Pro-single-token-entropy
Dataset Card for MMLU Pro with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
MMLU Pro dataset with single token response entropy metadata for Mistral 24B, Phi4, Phi4-mini, Qwen2.5 3B
Dataset Details
Dataset Description
Following up on the results from "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", we measure the response entopy for MMLU Pro dataset when the model is prompted to… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-single-token-entropy.MMLU-Pro-reasoning-score
Dataset Card for MMLU Pro with reasoning scores
MMLU Pro dataset with reasoning scores
Dataset Details
Dataset Description
As discovered in "When an LLM is apprehensive about its answers -- and when its uncertainty is justified", amount of reasoning required to answer a question (a.k.a. reasoning score) is a beter metric to estimate model uncertainty compared to more human-like level of education. Following the foot steps outlined in that paper, we ask a… See the full description on the dataset page: https://huggingface.co/datasets/LabARSS/MMLU-Pro-reasoning-score.turkish-grammar-mmlu
Turkish-Grammar-MMLU
This dataset, created by Turkish-DB, is a multiple-choice question-answering (QA) dataset covering Turkish grammar topics. It is designed to evaluate model performance on various Turkish grammar subjects, similar to the MMLU (Massive Multitask Language Understanding) benchmark.
Overview
Name: Turkish-Grammar-MMLU
Provider: Turkish-DB
Task: Multiple-Choice QA
Modality: Text
Format: CSV (also accessible via API in Parquet format)
Language: Turkish… See the full description on the dataset page: https://huggingface.co/datasets/turkish-db/turkish-grammar-mmlu.mmlu_reasoning100 samples from each topic of mmlu test. 57 * 100 = 5700 samples.
Added few shot examples generated by Llama-3.3-70B.
mmlu_trMobile-MMLU-Prommlu-reasoningtiny-re-MMLUMMLU_English_Choices_Aturkish_mmlu_with_reasoningMMLU_EN_USHello team,
I know OpenAI did a great job with the translations of the MMLU dataset into different languages, but I noticed the English CSV was missing. So, I took the liberty of unifying the files into the same structure and here I am sharing the complete MMLU dataset in English.
In the future, I plan to add a combined test in both Spanish and English, but for now, that's all. Thanks for checking it out!
Arabic-Cohere-include-base-44-mmlu-style
The Refined Arabic Cohere INCLUDE Base 44 Dataset as MMLU-Style
Dataset Summary
INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark spanning 44 languages that evaluates multilingual LLMs in the actual linguistic environments where they are deployed. The original dataset contains 22,637 4-option multiple-choice questions (MCQs) extracted from academic and professional exams, covering 57 topics, including regional knowledge.
When we reviewed the Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-Cohere-include-base-44-mmlu-style.global-mmlu-lite
Global MMLU Lite – Galician & Urdu
Machine-translated Galician and Urdu subsets of the Global MMLU Lite benchmark.
Dataset Description
Global MMLU Lite is a culturally-aware, multilingual evaluation benchmark for large language models, covering multiple-choice questions across many academic subjects. This repository contains Galician (gl) and Urdu (ur) translations. This dataset was translated using Google Machine Translate.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/Owos/global-mmlu-lite.MMLU-Phrasing-Benchmark
MMLU Phrasing Benchmark
This dataset is a phrasing variant of cais/mmlu, put together by Roscommon Systems to see whether the way a question is worded affects how accurately language models answer it.
Each of the 2,650 questions appears four ways: the original text from MMLU, a polite version, a formal academic version, and an angry/demanding version. The answer choices and correct answers are identical to the source dataset in all cases.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/RoscommonSystems/MMLU-Phrasing-Benchmark.twice_kr_financial_mmlu_cls
FinancialMMLU-CLS-ko
Multiple-choice questions, where a question and answer choices are provided to find the correct answer.
An open dataset generated and verified by GPT, based on financial public websites and Wikipedia.
Utilizing the open dataset allganize/financial-mmlu-ko (original source: public websites, Wikipedia).
llm-metric-MMLU-Prommlu-pro-enrichedmmlu_trainsetMMLU_English_Choices_C
