CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Idavidrein /gpqagated Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation… See the full description on the dataset page: https://huggingface.co/datasets/Idavidrein/gpqa.tabularquestion-answering1K<n<10K542 likes128k downloads1d agoHugging Face02fingertap /GPQA-Diamondtextn<1K16 likes5.4k downloads1y agoHugging Face03nmayorga7 /gpqa_diamondtabularn<1K0 likes5.4k downloads1y agoHugging Face04hendrydong /gpqa_diamond_mctextn<1K2 likes3.2k downloads2y agoHugging Face05hendrydong /gpqa_diamondtextn<1K10 likes2.8k downloads2y agoHugging Face06hendrydong /gpqa_main_mctextn<1K1 likes2.5k downloads2y agoHugging Face07hendrydong /gpqa_maintextn<1K1 likes2.5k downloads2y agoHugging Face08Wanfq /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.tabularquestion-answering1K<n<10K0 likes2.4k downloads2y agoHugging Face09ankner /gpqatabularn<1K0 likes2.2k downloads2y agoHugging Face10mkhalifa /gpqa-diamond-physicstextn<1K0 likes1.8k downloads2y agoHugging Face11zekeZZ /gpqa_alltext1K<n<10K0 likes1.5k downloads2y agoHugging Face12RaccoonOnion /gpqa-swaptabularquestion-answering1K<n<10K0 likes1.5k downloads1y agoHugging Face13PNYX /gpqa_subtask GPQA Subtask This is an splitted version of the GPQA dataset, where different domains and subdomains are in different files. Samples per subdomain in main: biology 78 physics 187 chemistry 183 Samples per subdomain in diamond: physics 86 chemistry 93 biology 19 Samples per subdomain in extended: biology 105 physics 227 chemistry 214 Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/gpqa_subtask.tabularquestion-answering1K<n<10K0 likes1.3k downloads1y agoHugging Face14lmarena-ai /PPE-GPQA-Best-of-K Overview This contains the GPQA correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from GPQA. This dataset is meant for benchmarking and evaluation, not for training. Paper Code License User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use set by the respective model providers. Citation @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward Models for… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/PPE-GPQA-Best-of-K.tabularn<1K1 likes1.2k downloads2y agoHugging Face15jeggers /gpqa_formattedgated Dataset Card for GPQA Formatted version of original GPQA dataset. This removes most columns and adds single columns options and answer to contain a list of the possible answers and the index of the correct one. GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy… See the full description on the dataset page: https://huggingface.co/datasets/jeggers/gpqa_formatted.textn<1K4 likes1.1k downloads2y agoHugging Face16hazyresearch /GPQA_GPT-4o Features id: Unique identifier for each problem question: Original question text options: List of 4 answer options, where first element is the correct answer prompts: List of 1000 prompts with randomized option orders for each problem. responses: List of 1000 model responses for each problem correct_bools: List of 1000 boolean values indicating if each response was correct split: Dataset split identifier ('diamond' or 'main') textn<1K0 likes903 downloads2y agoHugging Face17qfq /train_gpqa_update_testtext10K<n<100K0 likes803 downloads2y agoHugging Face18aradhye /gpqa_diamondtextn<1K0 likes656 downloads1y agoHugging Face19mlfoundations-dev /sci_question_exp__scp_116k__training_2k_for_GPQAtext100K<n<1M1 likes616 downloads2y agoHugging Face20ariaattarml /verified-reasoning-o1-gpqa-mmlu-pro Reasoning PRM Preference Dataset This dataset contains reasoning traces from multiple sources (GPQA Diamond and MMLU Pro), labeled with preference information based on correctness verification. Dataset Description Overview The dataset consists of reasoning problems and their solutions, where each example has been verified for correctness and labeled with a preference score. It combines data from two main sources: GPQA Diamond MMLU Pro Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ariaattarml/verified-reasoning-o1-gpqa-mmlu-pro.textn<1K2 likes510 downloads2y agoHugging Face21zekeZZ /gpqa_chemtextn<1K1 likes504 downloads2y agoHugging Face22AdnanElAssadi /Google-Translated_Turkish_GPQA_Datasettabular1K<n<10K0 likes485 downloads2y agoHugging Face23math-ai /gpqatextn<1K1 likes473 downloads2y agoHugging Face24zekeZZ /gpqa_biotextn<1K0 likes439 downloads2y agoHugging Face25zekeZZ /gpqa_physicstextn<1K0 likes433 downloads2y agoHugging Face26nichenshun /gpqa_diamondtextn<1K0 likes397 downloads2y agoHugging Face27ellamind /gpqa-multilingualgated GPQA Multilingual Multilingual translations of GPQA (Graduate-Level Google-Proof Q&A), a challenging multiple-choice benchmark requiring graduate-level expertise in biology, physics, and chemistry. Source: Idavidrein/gpqa (gpqa_main, 448 questions) Languages Config Language Examples ces Czech 448 dan Danish 448 deu German 448 fin Finnish 50 fra French 448 ita Italian 448 nld Dutch 448 pol Polish 448 spa Spanish 448 More to be added later.… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/gpqa-multilingual.textquestion-answering1K<n<10K0 likes389 downloads7mo agoHugging Face28Complementarity /gpqa-metadata-blind-answertabularn<1K0 likes387 downloads1mo agoHugging Face29hazyresearch /GPQA_GPT-4o-mini Features id: Unique identifier for each problem question: Original question text options: List of 4 answer options, where first element is the correct answer prompts: List of 1000 prompts with randomized option orders for each problem. responses: List of 1000 model responses for each problem correct_bools: List of 1000 boolean values indicating if each response was correct split: Dataset split identifier ('diamond' or 'main') textn<1K0 likes385 downloads2y agoHugging Face30shanchen /gpqa_diamond_mc_multilingualWhen Models Reason in Your Language: Controlling Thinking Trace Language Comes at the Cost of Accuracy https://arxiv.org/abs/2505.22888 Jirui Qi, Shan Chen, Zidi Xiong, Raquel Fernández, Danielle S. Bitterman, Arianna Bisazza Recent Large Reasoning Models (LRMs) with thinking traces have shown strong performance on English reasoning tasks. However, their ability to think in other languages is less studied. This capability is as important as answer accuracy for real world applications because… See the full description on the dataset page: https://huggingface.co/datasets/shanchen/gpqa_diamond_mc_multilingual.text1K<n<10K2 likes380 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.