CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Idavidrein /gpqagated Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.tabularquestion-answering1K<n<10K541 likes128k downloads21h agoHugging Face02nmayorga7 /gpqa_diamondtabularn<1K0 likes5.4k downloads1y agoHugging Face03Wanfq /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.tabularquestion-answering1K<n<10K0 likes2.4k downloads2y agoHugging Face04ankner /gpqatabularn<1K0 likes2.2k downloads2y agoHugging Face05RaccoonOnion /gpqa-swaptabularquestion-answering1K<n<10K0 likes1.5k downloads1y agoHugging Face06PNYX /gpqa_subtask GPQA Subtask This is an splitted version of the GPQA dataset, where different domains and subdomains are in different files. Samples per subdomain in main: biology 78 physics 187 chemistry 183 Samples per subdomain in diamond: physics 86 chemistry 93 biology 19 Samples per subdomain in extended: biology 105 physics 227 chemistry 214 Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/gpqa_subtask.tabularquestion-answering1K<n<10K0 likes1.3k downloads1y agoHugging Face07lmarena-ai /PPE-GPQA-Best-of-K Overview This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from GPQA. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward Models for… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-GPQA-Best-of-K.tabularn<1K1 likes1.2k downloads2y agoHugging Face08AdnanElAssadi /Google-Translated_Turkish_GPQA_Datasettabular1K<n<10K0 likes485 downloads2y agoHugging Face09Complementarity /gpqa-metadata-blind-answertabularn<1K0 likes387 downloads29d agoHugging Face10llamastack /gpqa_0shot_cottabular1K<n<10K0 likes350 downloads6mo agoHugging Face11JingzeShi /gpqa_diamondtabularn<1K0 likes232 downloads1y agoHugging Face12teddyyyy123 /gpqa_0shot_cottabular1K<n<10K0 likes222 downloads2y agoHugging Face13natong19 /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/natong19/gpqa.tabularquestion-answering1K<n<10K0 likes189 downloads9mo agoHugging Face14ivanvmoreno /gpqa-mttabular10K<n<100K0 likes158 downloads9mo agoHugging Face15jimim /gpqa_subset100tabularn<1K0 likes156 downloads2y agoHugging Face16nikhilchandak /gpqa-diamond-annotations GPQA Diamond Dataset This dataset contains filtered JSONL files of human annotations on question specificity, answer uniqueness, answer matching to the ground truth for different models for the GPQA Diamond dataset. The dataset was annotated by two human graders. It contains 198 (original size) * 2 = 396 rows as each rows is repeated twice (one for each human). A human grader given the question, actual answer and model response, has to answer whether the response matches the… See the full description on the dataset page: https://huggingface.co/datasets/nikhilchandak/gpqa-diamond-annotations.tabularn<1K1 likes148 downloads1y agoHugging Face17joanvelja /gpqa-open-ended GPQA Open-Ended An open-ended reformulation of the GPQA (Graduate-Level Google-Proof Questions) benchmark. All 546 questions have been converted from multiple-choice to free-response format, preserving the original difficulty and domain expertise requirements while removing the ability to eliminate answers or pattern-match against option structure. Why open-ended? MCQ benchmarks have a ceiling problem for scalable oversight research: a non-expert judge who cannot solve… See the full description on the dataset page: https://huggingface.co/datasets/joanvelja/gpqa-open-ended.tabularquestion-answering1K<n<10K0 likes146 downloads6mo agoHugging Face18stewy33 /acc_rd_s1-gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.tabularquestion-answering1K<n<10K0 likes142 downloads2y agoHugging Face19llamastack /gpqa_0shottabular1K<n<10K0 likes142 downloads1y agoHugging Face20johnsonafool /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/johnsonafool/gpqa.tabularquestion-answering1K<n<10K0 likes127 downloads1y agoHugging Face21mlfoundations-dev /sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912 mlfoundations-dev/sci_question_exp__scp_116k__training_2k_for_GPQA_eval_03-11-25_17-16-22_f912 Precomputed model outputs for evaluation. Evaluation Results GPQADiamond Average Accuracy: 26.94% ± 4.54% Number of Runs: 3 Run Accuracy Questions Solved Total Questions 1 19.70% 39 198 2 23.23% 46 198 3 37.88% 75 198 tabularn<1K1 likes116 downloads2y agoHugging Face22bdytx5 /gpqa_gpqa_main Forked Dataset This is a forked version of Idavidrein/gpqa (config: gpqa_main) Forked on: October 22, 2025 Modifications Added question_id column: Unique integer ID for each question across all splits (starting from 0) Original Dataset For the original dataset, please visit: https://huggingface.co/datasets/Idavidrein/gpqa tabularn<1K0 likes103 downloads11mo agoHugging Face23keenanpepper /gpqa-misleading-hintsThis dataset is an extension to Idavidrein/gpqa (GPQA dataset) that consists of one "Misleading Hint" for each question. The purpose of the hints is to lead the test taker down an invalid logical path that points them toward the wrong answer. In experiments with meta-llama/Llama-3.3-70B-Instruct , including the hints indeed decreases the performance very significantly. The hints were generated by Claude 3.5 Sonnet in a multi-step process involving generating three candidate hints and choosing… See the full description on the dataset page: https://huggingface.co/datasets/keenanpepper/gpqa-misleading-hints.tabularn<1K0 likes102 downloads2y agoHugging Face24kth8 /Qwen3.5-27B-AWQ-4bit-GPQA-Diamond-benchmarkBenchmark of cyankiwi/Qwen3.5-27B-AWQ-4bit against fingertap/GPQA-Diamond dataset. Accuracy: 76.3% with Python tool. Metric Value Correct 151 Incorrect 46 Errors 1 Total samples 198 Python tool calls 225 Total completion tokens 659,879 Raw stats: { "accuracy": 0.763, "correct": 151, "incorrect": 46, "error": 1, "total": 198, "python_tool_calls": 225, "completion_tokens": 659879 } tabularn<1K0 likes93 downloads6mo agoHugging Face25alperengozeten /gpqa_with_summariestabularn<1K0 likes92 downloads2y agoHugging Face26hbXNov /gpqa_diamond_64tabularn<1K0 likes90 downloads2y agoHugging Face27HAERAE-HUB /hret_agent_idavidrein_gpqa_diamond_translatedtabularn<1K0 likes80 downloads2y agoHugging Face28nishadsinghi /gpqa_diamond_64tabularn<1K0 likes78 downloads2y agoHugging Face29mlfoundations-dev /GPQADiamond_evalchemytabular1K<n<10K0 likes78 downloads2y agoHugging Face30XiangPan /gpqa-mctabularn<1K0 likes78 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.