mbart
Datasets
All datasets matching “mbart”conceptual-12m-mbart-50-multilingualaffilgood-mBART-rorLeo__mbart-large-cc25__1645802644sciq-ja-mbartm2m
Dataset Card for "sciq-ja-mbartm2m"
Dataset Description
This is the Japanese Translation version of sciq.
The translator used in it was facebook/mbart-large-50-many-to-many-mmt.
License
The same as the original sciq (cc-by-nc-3.0).
synQASynQA is a Reading Comprehension dataset created in the work "Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation" (https://aclanthology.org/2021.emnlp-main.696/).
It consists of 314,811 synthetically generated questions on the passages in the SQuAD v1.1 (https://arxiv.org/abs/1606.05250) training set.
In this work, we use a synthetic adversarial data generation to make QA models more robust to human adversaries. We develop a data generation pipeline that selects source passages, identifies candidate answers, generates questions, then finally filters or re-labels them to improve quality. Using this approach, we amplify a smaller human-written adversarial dataset to a much larger set of synthetic question-answer pairs. By incorporating our synthetic data, we improve the state-of-the-art on the AdversarialQA (https://adversarialqa.github.io/) dataset by 3.7F1 and improve model generalisation on nine of the twelve MRQA datasets. We further conduct a novel human-in-the-loop evaluation to show that our models are considerably more robust to new human-written adversarial examples: crowdworkers can fool our model only 8.8% of the time on average, compared to 17.6% for a model trained without synthetic data.
For full details on how the dataset was created, kindly refer to the paper.piqa-ja-mbartm2m
Dataset Card for "piqa-ja-mbartm2m"
Dataset Description
This is the Japanese Translation version of piqa.
The translator used in it was facebook/mbart-large-50-many-to-many-mmt.
License
The same as the original piqa.
