datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.gpqa_diamondgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.gpqagpqa-swapgpqa_subtask
GPQA Subtask
This is an splitted version of the GPQA dataset, where different domains and subdomains are in different files.
Samples per subdomain in main:
biology 78
physics 187
chemistry 183
Samples per subdomain in diamond:
physics 86
chemistry 93
biology 19
Samples per subdomain in extended:
biology 105
physics 227
chemistry 214
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/gpqa_subtask.PPE-GPQA-Best-of-K
Overview
This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from GPQA.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
License
User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers.
Citation
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-GPQA-Best-of-K.Google-Translated_Turkish_GPQA_Datasetgpqa-metadata-blind-answergpqa_0shot_cotgpqa_diamondgpqa_0shot_cotgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.gpqa-mtgpqa_subset100gpqa-diamond-annotations
GPQA Diamond Dataset
This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset.
The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human).
A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.gpqa-open-ended
GPQA Open-Ended
An open-ended reformulation of the GPQA (Graduate-Level Google-Proof Questions) benchmark. All 546 questions have been converted from multiple-choice to free-response format, preserving the original difficulty and domain expertise requirements while removing the ability to eliminate answers or pattern-match against option structure.
Why open-ended?
MCQ benchmarks have a ceiling problem for scalable oversight research: a non-expert judge who cannot solve… See the full description on the dataset page: https://huggingface.co/datasets/joanvelja/gpqa-open-ended.acc_rd_s1-gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.gpqa_0shotgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/johnsonafool/gpqa.sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
mlfoundations-dev/sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 26.94% ± 4.54%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
19.70%
39
198
2
23.23%
46
198
3
37.88%
75
198
gpqa_gpqa_main
Forked Dataset
This is a forked version of Idavidrein/gpqa (config: gpqa_main)
Forked on: October 22, 2025
Modifications
Added question_id column: Unique integer ID for each question across all splits (starting from 0)
Original Dataset
For the original dataset, please visit: https://huggingface.co/datasets/Idavidrein/gpqa
gpqa-misleading-hintsThis dataset is an extension to Idavidrein/gpqa (GPQA dataset) that consists of one "Misleading Hint" for each question.
The purpose of the hints is to lead the test taker down an invalid logical path that points them toward the wrong answer. In experiments with meta-llama/Llama-3.3-70B-Instruct , including the hints indeed decreases the performance very significantly.
The hints were generated by Claude 3.5 Sonnet in a multi-step process involving generating three candidate hints and choosing… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/gpqa-misleading-hints.Qwen3.5-27B-AWQ-4bit-GPQA-Diamond-benchmarkBenchmark of cyankiwi/Qwen3.5-27B-AWQ-4bit against fingertap/GPQA-Diamond dataset.
Accuracy: 76.3% with Python tool.
Metric
Value
Correct
151
Incorrect
46
Errors
1
Total samples
198
Python tool calls
225
Total completion tokens
659,879
Raw stats:
{
"accuracy": 0.763,
"correct": 151,
"incorrect": 46,
"error": 1,
"total": 198,
"python_tool_calls": 225,
"completion_tokens": 659879
}
gpqa_with_summariesgpqa_diamond_64hret_agent_idavidrein_gpqa_diamond_translatedgpqa_diamond_64GPQADiamond_evalchemygpqa-mc
