datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.GPQA-Diamondgpqa_diamondgpqa_diamond_mcgpqa_diamondgpqa_main_mcgpqa_maingpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.gpqagpqa-diamond-physicsgpqa_allgpqa-swapgpqa_subtask
GPQA Subtask
This is an splitted version of the GPQA dataset, where different domains and subdomains are in different files.
Samples per subdomain in main:
biology 78
physics 187
chemistry 183
Samples per subdomain in diamond:
physics 86
chemistry 93
biology 19
Samples per subdomain in extended:
biology 105
physics 227
chemistry 214
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/gpqa_subtask.PPE-GPQA-Best-of-K
Overview
This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from GPQA.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
License
User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers.
Citation
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-GPQA-Best-of-K.gpqa_formatted
Dataset Card for GPQA
Formatted version of original GPQA dataset. This removes most columns and adds single columns options and answer to contain a list of the possible answers and the index of the correct one.
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/gpqa_formatted.GPQA_GPT-4o
Features
id: Unique identifier for each problem
question: Original question text
options: List of 4 answer options, where first element is the correct answer
prompts: List of 1000 prompts with randomized option orders for each problem.
responses: List of 1000 model responses for each problem
correct_bools: List of 1000 boolean values indicating if each response was correct
split: Dataset split identifier ('diamond' or 'main')
train_gpqa_update_testgpqa_diamondsci_question_exp__scp_116k__training_2k_for_GPQAverified-reasoning-o1-gpqa-mmlu-pro
Reasoning PRM Preference Dataset
This dataset contains reasoning traces from multiple sources (GPQA Diamond and MMLU Pro), labeled with preference information based on correctness verification.
Dataset Description
Overview
The dataset consists of reasoning problems and their solutions, where each example has been verified for correctness and labeled with a preference score. It combines data from two main sources:
GPQA Diamond
MMLU Pro
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ariaattarml/verified-reasoning-o1-gpqa-mmlu-pro.gpqa_chemGoogle-Translated_Turkish_GPQA_Datasetgpqagpqa_biogpqa_physicsgpqa_diamondgpqa-multilingual
GPQA Multilingual
Multilingual translations of GPQA (Graduate-Level Google-Proof Q&A), a challenging multiple-choice benchmark requiring graduate-level expertise in biology, physics, and chemistry.
Source: Idavidrein/gpqa (gpqa_main, 448 questions)
Languages
Config
Language
Examples
ces
Czech
448
dan
Danish
448
deu
German
448
fin
Finnish
50
fra
French
448
ita
Italian
448
nld
Dutch
448
pol
Polish
448
spa
Spanish
448
More to be added later.… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/gpqa-multilingual.gpqa-metadata-blind-answerGPQA_GPT-4o-mini
Features
id: Unique identifier for each problem
question: Original question text
options: List of 4 answer options, where first element is the correct answer
prompts: List of 1000 prompts with randomized option orders for each problem.
responses: List of 1000 model responses for each problem
correct_bools: List of 1000 boolean values indicating if each response was correct
split: Dataset split identifier ('diamond' or 'main')
gpqa_diamond_mc_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy
https://arxiv.org/abs/2505.22888
Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza
Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/gpqa_diamond_mc_multilingual.
