datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.gpqa_diamondGPQA-Diamondgpqa_diamond_mcgpqa_diamondgpqa_main_mcgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.gpqa_maingpqagpqa-diamond-physicsgpqa_allgpqa-swapgpqa_subtask
GPQA Subtask
This is an splitted version of the GPQA dataset, where different domains and subdomains are in different files.
Samples per subdomain in main:
biology 78
physics 187
chemistry 183
Samples per subdomain in diamond:
physics 86
chemistry 93
biology 19
Samples per subdomain in extended:
biology 105
physics 227
chemistry 214
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/gpqa_subtask.PPE-GPQA-Best-of-K
Overview
This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations.
The prompts are sampled from GPQA.
This dataset is meant for benchmarking and evaluation, not for training.
Paper
Code
License
User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers.
Citation
@misc{frick2024evaluaterewardmodelsrlhf,
title={How to Evaluate Reward Models for… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-GPQA-Best-of-K.gpqa_formatted
Dataset Card for GPQA
Formatted version of original GPQA dataset. This removes most columns and adds single columns options and answer to contain a list of the possible answers and the index of the correct one.
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/gpqa_formatted.gpqasynthetic-gpqa-oct4GPQA_GPT-4o
Features
id: Unique identifier for each problem
question: Original question text
options: List of 4 answer options, where first element is the correct answer
prompts: List of 1000 prompts with randomized option orders for each problem.
responses: List of 1000 model responses for each problem
correct_bools: List of 1000 boolean values indicating if each response was correct
split: Dataset split identifier ('diamond' or 'main')
train_gpqa_update_testloom-benchmark-gpqagpqa_diamondsci_question_exp__scp_116k__training_2k_for_GPQAverified-reasoning-o1-gpqa-mmlu-pro
Reasoning PRM Preference Dataset
This dataset contains reasoning traces from multiple sources (GPQA Diamond and MMLU Pro), labeled with preference information based on correctness verification.
Dataset Description
Overview
The dataset consists of reasoning problems and their solutions, where each example has been verified for correctness and labeled with a preference score. It combines data from two main sources:
GPQA Diamond
MMLU Pro
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ariaattarml/verified-reasoning-o1-gpqa-mmlu-pro.gpqa_chemGoogle-Translated_Turkish_GPQA_Datasetgpqagpqa_biogpqa_physicsloom-benchmark-gpqa-diamondGPQA_GPT-4o-mini
Features
id: Unique identifier for each problem
question: Original question text
options: List of 4 answer options, where first element is the correct answer
prompts: List of 1000 prompts with randomized option orders for each problem.
responses: List of 1000 model responses for each problem
correct_bools: List of 1000 boolean values indicating if each response was correct
split: Dataset split identifier ('diamond' or 'main')
