datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BoolQ_robustness
Dataset Card for "BoolQ-robustness"
Dataset Summary
BoolQ-robustness is an expanded version of the BoolQ dataset (https://arxiv.org/abs/1905.10044) but with perturbations of the original input questions and passages.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
boolq_robustness
Size of downloaded dataset file: 21.8 MB
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/BoolQ_robustness.PopQA_robustness
Dataset Card for "PopQA-robustness"
Dataset Summary
PopQS-robustness is an expanded version of the PopQA dataset (https://aclanthology.org/2023.acl-long.546/) but with perturbations of the original input questions.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
popqa_robustness
Size of downloaded dataset file: 26.4 MB
Data Fields
boolq_robustness… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/PopQA_robustness.identity_group_abuse_robustness
Dataset Card for "identity_group_abuse-robustness"
Dataset Summary
identity_group_abuse-robustness is an expanded version of the identity group abuse dataset (https://aclanthology.org/2022.naacl-main.410/) but with perturbations of the original input questions and passages.
It is intended for use as a benchmark for evaluating model robustness on question-answering to these perturbations.
Data Instances
identity_group_abuse-robustness
Size of… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/identity_group_abuse_robustness.
