datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wikipedia_contradict_benchmark
Wikipedia contradict benchmark
Wikipedia contradict benchmark is a dataset consisting of 253 high-quality, human-annotated instances designed to assess LLM performance when augmented with retrieved passages containing real-world knowledge conflicts. The dataset was created intentionally with that task in mind, focusing on a benchmark consisting of high-quality, human-annotated instances.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/Wikipedia_contradict_benchmark.SocialStigmaQA
SocialStigmaQA Dataset Card
Current datasets for unwanted social bias auditing are limited to studying protected demographic features such as race and gender.
In this dataset, we introduce a dataset that is meant to capture the amplification of social bias, via stigmas, in generative language models.
Taking inspiration from social science research, we start with a documented list of 93 US-centric stigmas and curate a question-answering (QA) dataset which involves simple social… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA.BPC
BPC: A Benchmark Dataset for Causal Business Process Reasoning
Dataset Card for BPC
Dataset Summary
Abstract. Large Language Models (LLMs) are increasingly used for boosting organizational efficiency and automating tasks.
While not originally designed for complex cognitive processes, recent efforts have further extended to employ LLMs in activities such as reasoning, planning,
and decision-making. In business processes, such abilities could be invaluable for… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BPC.SocialStigmaQA-JA
SocialStigmaQA-JA Dataset Card
It is crucial to test the social bias of large language models.
SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models.
Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations.
Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.BoolQ_robustness
Dataset Card for "BoolQ-robustness"
Dataset Summary
BoolQ-robustness is an expanded version of the BoolQ dataset (https://arxiv.org/abs/1905.10044) but with perturbations of the original input questions and passages.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
boolq_robustness
Size of downloaded dataset file: 21.8 MB
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BoolQ_robustness.PopQA_robustness
Dataset Card for "PopQA-robustness"
Dataset Summary
PopQS-robustness is an expanded version of the PopQA dataset (https://aclanthology.org/2023.acl-long.546/) but with perturbations of the original input questions.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
popqa_robustness
Size of downloaded dataset file: 26.4 MB
Data Fields
boolq_robustness… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/PopQA_robustness.AttaQ-JA
AttaQ-JA Dataset Card
AttaQ red teaming dataset was designed to evaluate Large Language Models (LLMs) by assessing their tendency to generate harmful or undesirable responses, which consists of 1402 carefully crafted adversarial questions.
This AttaQ-JA dataset is a Japanese version of AttaQ, created by translating manually and carefully.
Disclaimer:
The data contains offensive and upsetting content by nature, therefore it may not be easy to read. Please read them in… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AttaQ-JA.identity_group_abuse_robustness
Dataset Card for "identity_group_abuse-robustness"
Dataset Summary
identity_group_abuse-robustness is an expanded version of the identity group abuse dataset (https://aclanthology.org/2022.naacl-main.410/) but with perturbations of the original input questions and passages.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
identity_group_abuse-robustness
Size of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/identity_group_abuse_robustness.
