CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes793 downloads2y agoHugging Face02GenData-Research /scientific-verification Scientific Verification Benchmark: NMC Cathodes Dataset summary The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.tabularquestion-answering1K<n<10K0 likes241 downloads9d agoHugging Face03macpaw-research /UiPad UiPad - UI Parsing and Accessibility Dataset 📌 Dataset status: stable release. UiPad was built for the IASA Champ 2024 Challenge and is a complete, fixed research artifact. No further updates are planned. Curated by: MacPaw Way Ltd. Language(s): Mostly EN, UA License: MIT Overview UiPad is a dataset created for the IASA Champ 2024 Challenge, focusing on the accessibility and interface understanding of MacOS applications. With growing interest in AI-driven user interface… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/UiPad.imagequestion-answering1K<n<10K16 likes200 downloads1mo agoHugging Face04ciol-research /multilevel-legal-reasoning Legal Reasoning Dataset with Multilevel Human and Model-Annotated Explanations Prepared by Mst Rafia Islam, Umong Sain, Azmine Toushik Wasi Prepared as a part of Reasoning Datasets Competition by Bespoke Labs, Hugging Face, and Together.ai. 🧭 Purpose and Scope The Legal Reasoning Dataset aims to support the evaluation and training of legal reasoning systems, particularly in multilingual or jurisdiction-agnostic contexts. It focuses on international acts and treaties… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/multilevel-legal-reasoning.tabulartext-generationn<1K7 likes65 downloads1y agoHugging Face05ibm-research /SocialStigmaQA SocialStigmaQA Dataset Card Current datasets for unwanted social bias auditing are limited to studying protected demographic features such as race and gender. In this dataset, we introduce a dataset that is meant to capture the amplification of social bias, via stigmas, in generative language models. Taking inspiration from social science research, we start with a documented list of 93 US-centric stigmas and curate a question-answering (QA) dataset which involves simple social… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA.textquestion-answering10K<n<100K7 likes63 downloads2y agoHugging Face06ibm-research /BPC BPC: A Benchmark Dataset for Causal Business Process Reasoning Dataset Card for BPC Dataset Summary Abstract. Large Language Models (LLMs) are increasingly used for boosting organizational efficiency and automating tasks. While not originally designed for complex cognitive processes, recent efforts have further extended to employ LLMs in activities such as reasoning, planning, and decision-making. In business processes, such abilities could be invaluable for… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BPC.textquestion-answering1K<n<10K3 likes56 downloads2y agoHugging Face07ibm-research /SocialStigmaQA-JA SocialStigmaQA-JA Dataset Card It is crucial to test the social bias of large language models. SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models. Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations. Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.tabularquestion-answering10K<n<100K4 likes46 downloads2y agoHugging Face08ibm-research /BoolQ_robustness Dataset Card for "BoolQ-robustness" Dataset Summary BoolQ-robustness is an expanded version of the BoolQ dataset (https://arxiv.org/abs/1905.10044) but with perturbations of the original input questions and passages. It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations. Data Instances boolq_robustness Size of downloaded dataset file: 21.8 MB Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BoolQ_robustness.tabularquestion-answering10K<n<100K0 likes37 downloads2y agoHugging Face09ciol-research /BengaliMoralBench BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture Accepted at ACM FAccT 2026 · Montreal, QC, Canada · June 25–28, 2026 View in arXiv: https://arxiv.org/abs/2511.03180 View website: https://ciol-researchlab.github.io/works/BengaliMoralBench/ 📋 Overview BengaliMoralBench is the first large-scale, culturally grounded ethics benchmark for evaluating moral reasoning in Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/ciol-research/BengaliMoralBench.tabularquestion-answering1K<n<10K0 likes34 downloads4mo agoHugging Face10ibm-research /PopQA_robustness Dataset Card for "PopQA-robustness" Dataset Summary PopQS-robustness is an expanded version of the PopQA dataset (https://aclanthology.org/2023.acl-long.546/) but with perturbations of the original input questions. It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations. Data Instances popqa_robustness Size of downloaded dataset file: 26.4 MB Data Fields boolq_robustness… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/PopQA_robustness.tabularquestion-answering100K<n<1M0 likes32 downloads2y agoHugging Face11ibm-research /AttaQ-JA AttaQ-JA Dataset Card AttaQ red teaming dataset was designed to evaluate Large Language Models (LLMs) by assessing their tendency to generate harmful or undesirable responses, which consists of 1402 carefully crafted adversarial questions. This AttaQ-JA dataset is a Japanese version of AttaQ, created by translating manually and carefully. Disclaimer: The data contains offensive and upsetting content by nature, therefore it may not be easy to read. Please read them in… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AttaQ-JA.textquestion-answering1K<n<10K2 likes26 downloads2y agoHugging Face12ibm-research /identity_group_abuse_robustness Dataset Card for "identity_group_abuse-robustness" Dataset Summary identity_group_abuse-robustness is an expanded version of the identity group abuse dataset (https://aclanthology.org/2022.naacl-main.410/) but with perturbations of the original input questions and passages. It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations. Data Instances identity_group_abuse-robustness Size of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/identity_group_abuse_robustness.tabularquestion-answering10K<n<100K2 likes23 downloads2y agoHugging Face13asahi-research /newsqgated 時事情報に関する日本語QAデータセット『ニュースQ』 ニュースQ紹介ページ 利用規約はこちら 個人情報の取り扱い:利用申込の際にお預かりした個人情報(お名前、所属、利用目的、メールアドレス)は、下記の目的で利用し、弊社の個人情報保護方針に従って取り扱います。 本ツールの使用状況の確認 本人の所属が正しく申請されているかの確認 本ツールをご使用いただくために必要なご連絡(アップデートのご連絡等) 本ツールを使用した感想等を調査するためのご連絡 textquestion-answeringn<1K1 likes7 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.