datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
acp_bench
ACP Bench
🏠 Homepage •
📄 Paper •
📄 Paper
ACPBench is a benchmark dataset designed to evaluate the reasoning capabilities of large language models (LLMs) in the context of Action, Change, and Planning. It spans 13 diverse domains:
Blocksworld
Logistics
Grippers
Grid
Ferry
FloorTile
Rovers
VisitAll
Depot
Goldminer
Satellite
Swap
Alfworld
Task Types in ACPBench
ACPBench includes the following 8 reasoning tasks:
Action Applicability (app)… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/acp_bench.SocialStigmaQA-JA
SocialStigmaQA-JA Dataset Card
It is crucial to test the social bias of large language models.
SocialStigmaQA dataset is meant to capture the amplification of social bias, via stigmas, in generative language models.
Taking inspiration from social science research, the dataset is constructed from a documented list of 93 US-centric stigmas and a hand-curated question-answering (QA) templates which involves social situations.
Here, we introduce SocialStigmaQA-JA, a Japanese version of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SocialStigmaQA-JA.BoolQ_robustness
Dataset Card for "BoolQ-robustness"
Dataset Summary
BoolQ-robustness is an expanded version of the BoolQ dataset (https://arxiv.org/abs/1905.10044) but with perturbations of the original input questions and passages.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
boolq_robustness
Size of downloaded dataset file: 21.8 MB
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BoolQ_robustness.PopQA_robustness
Dataset Card for "PopQA-robustness"
Dataset Summary
PopQS-robustness is an expanded version of the PopQA dataset (https://aclanthology.org/2023.acl-long.546/) but with perturbations of the original input questions.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
popqa_robustness
Size of downloaded dataset file: 26.4 MB
Data Fields
boolq_robustness… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/PopQA_robustness.identity_group_abuse_robustness
Dataset Card for "identity_group_abuse-robustness"
Dataset Summary
identity_group_abuse-robustness is an expanded version of the identity group abuse dataset (https://aclanthology.org/2022.naacl-main.410/) but with perturbations of the original input questions and passages.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
identity_group_abuse-robustness
Size of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/identity_group_abuse_robustness.
