datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.gpqa_diamondgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.gpqa-swapgpqa_diamondgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.acc_rd_s1-gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/johnsonafool/gpqa.GPQA-Test1GPQA-Anno2gpqa-diamond-test2gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/gpqa.gpqako-gpqa
ko-gpqa
ko-gpqa is a Korean-translated version of the GPQA (Graduate-Level Google‑Proof Q&A) benchmark dataset, which consists of high-difficulty science questions. Introduced in this paper, GPQA is designed to go beyond simple fact retrieval and instead test an AI system’s ability to perform deep understanding and logical reasoning. It is particularly useful for evaluating true comprehension and inference capabilities in language models.
The Korean translation was performed using… See the full description on the dataset page: https://huggingface.co/datasets/davidkim205/ko-gpqa.gpqa-extended_trThis dataset is the Turkish translation of GPQA dataset(extended), which is one of the most-known benchmarks to evaluate LLMs performance on graduate level science questions.
You can find the paper of GPQA below.
https://arxiv.org/abs/2311.12022
Translation Methodology:
A subset of the columns of every row translated seperately using "google/gemma-3-27b-it" which has a good capacity for both the languages English and Turkish and performed good in scientific translation.
The following prompt… See the full description on the dataset page: https://huggingface.co/datasets/ytu-ce-cosmos/gpqa-extended_tr.gpqa-swap2ru_gpqa_diamond
Карточка датасета GPQA Diamond (перевод на русский язык)
Этот датасет представляет собой перевод на русский язык оригинального набора данных. GPQA — это набор вопросов и ответов с несколькими вариантами ответов. Полученные задания достаточно сложные и составленны и проверенны экспертами по биологии, физике и химии.
Здесь только diamond часть всего датасета - 200 наиболее сложных задач уровня PhD.
Описание
Датасет содержит 200 вопросов по биологии, физике и химии.… See the full description on the dataset page: https://huggingface.co/datasets/AvitoTech/ru_gpqa_diamond.gpqa-extended_TRcombined_gpqa_llama8b_oracle_difficultygpqa-diamond-cue-long-8192-mtcombined_gpqa_llama70b_oracle_difficultygpqa-diamond-base-summary-8192-mt
