CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibm-research /duorc Dataset Card for duorc Dataset Summary The DuoRC dataset is an English language dataset of questions and answers gathered from crowdsourced AMT workers on Wikipedia and IMDb movie plots. The workers were given freedom to pick answer from the plots or synthesize their own answers. It contains two sub-datasets - SelfRC and ParaphraseRC. SelfRC dataset is built on Wikipedia movie plots solely. ParaphraseRC has questions written from Wikipedia movie plots and the answers are… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/duorc.textquestion-answering100K<n<1M34 likes3.4k downloads3y agoHugging Face02ibm-research /acp_bench ACP Bench 🏠 Homepage • 📄 Paper • 📄 Paper ACPBench is a benchmark dataset designed to evaluate the reasoning capabilities of large language models (LLMs) in the context of Action, Change, and Planning. It spans 13 diverse domains: Blocksworld Logistics Grippers Grid Ferry FloorTile Rovers VisitAll Depot Goldminer Satellite Swap Alfworld Task Types in ACPBench ACPBench includes the following 8 reasoning tasks: Action Applicability (app)… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/acp_bench.tabularquestion-answering1K<n<10K13 likes2.6k downloads8mo agoHugging Face03ibm-research /VAKRA 🔷 VAKRA: A Benchmark for Evaluating Multi-Hop, Multi-Source Tool-Calling Capabilities in AI Agents VAKRA (eValuating API and Knowledge Retrieval Agents using multi-hop, multi-source dialogues) is a tool-grounded, executable benchmark designed to evaluate how well AI agents reason end-to-end in enterprise-like settings. Rather than testing isolated skills, VARKA measures compositional reasoning across APIs and documents, using full execution traces to assess whether agents can… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAKRA.textquestion-answering1K<n<10K46 likes1.6k downloads12d agoHugging Face04ibm-research /AssetOpsBench AssetOpsBench AssetOpsBench is a specialized benchmark designed for evaluating Large Language Models (LLMs) and Multi-Agent systems in industrial operations. It focuses on the intersection of sensor data interpretation, maintenance logic, and Prognostics and Health Management (PHM). The benchmark enables researchers to test how effectively AI agents can manage complex industrial assets, such as compressors and hydraulic pumps, by applying rule-based logic and diagnostic… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AssetOpsBench.textquestion-answeringn<1K46 likes809 downloads4mo agoHugging Face05ibm-research /Wikipedia_contradict_benchmark Wikipedia contradict benchmark Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.textquestion-answeringn<1K28 likes796 downloads2y agoHugging Face06ibm-research /FailureSensorIQ FailureSensorIQ Dataset FailureSensorIQ is a Multi-Choice QA (MCQA) dataset that explores the relationships between sensors and failure modes for 10 industrial assets. |Github | 🏆Leaderboard | 📖Paper | Dataset Summary FailureSensorIQ is a Multi-Choice QA (MCQA) dataset that explores the relationships between sensors and failure modes for 10 industrial assets. By only leveraging the information found in ISO documents, we developed a data generation pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/FailureSensorIQ.textquestion-answering1K<n<10K9 likes472 downloads7mo agoHugging Face07ibm-research /SQL-API-Bench Dataset Card for Dataset Name This dataset contains QA that requires DB and API access at the same time. It is composed of two new benchmarks consisting of questions whose answers require a combination of database and API calls, both of which are augmentations of the popular Spider dataset and benchmark. Benchmark I replaces a fraction of the real Spider database tables with equivalents that are executed via APIs. This allows us to directly test the mechanism by which database and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SQL-API-Bench.textquestion-answering1K<n<10K5 likes295 downloads1y agoHugging Face08ibm-research /nc-bench Dataset Card for Natural Conversation Benchmark (NC-Bench) Dataset Summary The Natural Conversation Benchmarks (NC-Bench) aim to answer the question: How well can generative AI converse like humans do? In other words, the benchmarks begin to measure the general conversational competence of large language models (LLMs). They do this by testing models' ability to generate an appropriate type of conversational action, or dialogue act, in response to a particular sequence of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/nc-bench.textquestion-answeringn<1K1 likes75 downloads8mo agoHugging Face09ibm-research /SocialStigmaQA SocialStigmaQA Dataset Card Current datasets for unwanted social bias auditing are limited to studying protected demographic features such as race and gender. In this dataset, we introduce a dataset that is meant to capture the amplification of social bias, via stigmas, in generative language models. Taking inspiration from social science research, we start with a documented list of 93 US-centric stigmas and curate a question-answering (QA) dataset which involves simple social… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA.textquestion-answering10K<n<100K7 likes65 downloads2y agoHugging Face10ibm-research /BPC BPC: A Benchmark Dataset for Causal Business Process Reasoning Dataset Card for BPC Dataset Summary Abstract. Large Language Models (LLMs) are increasingly used for boosting organizational efficiency and automating tasks. While not originally designed for complex cognitive processes, recent efforts have further extended to employ LLMs in activities such as reasoning, planning, and decision-making. In business processes, such abilities could be invaluable for… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BPC.textquestion-answering1K<n<10K3 likes56 downloads2y agoHugging Face11ibm-research /SocialStigmaQA-JA SocialStigmaQA-JA Dataset Card It is crucial to test the social bias of large language models. SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models. Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations. Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.tabularquestion-answering10K<n<100K4 likes47 downloads2y agoHugging Face12ibm-research /BoolQ_robustness Dataset Card for "BoolQ-robustness" Dataset Summary BoolQ-robustness is an expanded version of the BoolQ dataset (https://arxiv.org/abs/1905.10044) but with perturbations of the original input questions and passages. It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations. Data Instances boolq_robustness Size of downloaded dataset file: 21.8 MB Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BoolQ_robustness.tabularquestion-answering10K<n<100K0 likes38 downloads2y agoHugging Face13ibm-research /PopQA_robustness Dataset Card for "PopQA-robustness" Dataset Summary PopQS-robustness is an expanded version of the PopQA dataset (https://aclanthology.org/2023.acl-long.546/) but with perturbations of the original input questions. It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations. Data Instances popqa_robustness Size of downloaded dataset file: 26.4 MB Data Fields boolq_robustness… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/PopQA_robustness.tabularquestion-answering100K<n<1M0 likes33 downloads2y agoHugging Face14ibm-research /AttaQ-JA AttaQ-JA Dataset Card AttaQ red teaming dataset was designed to evaluate Large Language Models (LLMs) by assessing their tendency to generate harmful or undesirable responses, which consists of 1402 carefully crafted adversarial questions. This AttaQ-JA dataset is a Japanese version of AttaQ, created by translating manually and carefully. Disclaimer: The data contains offensive and upsetting content by nature, therefore it may not be easy to read. Please read them in… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AttaQ-JA.textquestion-answering1K<n<10K2 likes26 downloads2y agoHugging Face15ibm-research /identity_group_abuse_robustness Dataset Card for "identity_group_abuse-robustness" Dataset Summary identity_group_abuse-robustness is an expanded version of the identity group abuse dataset (https://aclanthology.org/2022.naacl-main.410/) but with perturbations of the original input questions and passages. It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations. Data Instances identity_group_abuse-robustness Size of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/identity_group_abuse_robustness.tabularquestion-answering10K<n<100K2 likes23 downloads2y agoHugging Face16ibm-research /POBS Preference, Opinion, and Belief Survey (POBS) POBS is a dataset of survey questions designed to uncover preferences, opinions, and beliefs on societal issues. Each row represents a question with its topic, options, and polarity. Columns: topic: Question topic category: Category question_id: Unique question ID question: Survey question text options: List of possible answers options_polarity: Numeric polarity for each option (where applicable) POBS: Preference, Opinion… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/POBS.textquestion-answeringn<1K0 likes21 downloads1y agoHugging Face17mahimairaja /ibm-hls-burn-vectorizedimagequestion-answering1K<n<10K0 likes19 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.