datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual-12m-mbart-50-multilingualaffilgood-mBART-rorLeo__mbart-large-cc25__1645802644sciq-ja-mbartm2m
Dataset Card for "sciq-ja-mbartm2m"
Dataset Description
This is the Japanese Translation version of sciq.
The translator used in it was facebook/mbart-large-50-many-to-many-mmt.
License
The same as the original sciq (cc-by-nc-3.0).
synQASynQA is a Reading Comprehension dataset created in the work "Improving Question Answering Model Robustness with Synthetic Adversarial Data Generation" (https://aclanthology.org/2021.emnlp-main.696/).
It consists of 314,811 synthetically generated questions on the passages in the SQuAD v1.1 (https://arxiv.org/abs/1606.05250) training set.
In this work, we use a synthetic adversarial data generation to make QA models more robust to human adversaries. We develop a data generation pipeline that selects source passages, identifies candidate answers, generates questions, then finally filters or re-labels them to improve quality. Using this approach, we amplify a smaller human-written adversarial dataset to a much larger set of synthetic question-answer pairs. By incorporating our synthetic data, we improve the state-of-the-art on the AdversarialQA (https://adversarialqa.github.io/) dataset by 3.7F1 and improve model generalisation on nine of the twelve MRQA datasets. We further conduct a novel human-in-the-loop evaluation to show that our models are considerably more robust to new human-written adversarial examples: crowdworkers can fool our model only 8.8% of the time on average, compared to 17.6% for a model trained without synthetic data.
For full details on how the dataset was created, kindly refer to the paper.piqa-ja-mbartm2m
Dataset Card for "piqa-ja-mbartm2m"
Dataset Description
This is the Japanese Translation version of piqa.
The translator used in it was facebook/mbart-large-50-many-to-many-mmt.
License
The same as the original piqa.
autotrain-data-mbart-finetune-hindi
AutoTrain Dataset for project: mbart-finetune-hindi
Dataset Description
This dataset has been automatically processed by AutoTrain for project mbart-finetune-hindi.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": "\u092e\u0928 \u0915\u0940 \u0917\u0939\u0930\u093e\u0907\u092f\u094b\u0902 \u092e\u0947\u0902… See the full description on the dataset page: https://huggingface.co/datasets/viditsorg/autotrain-data-mbart-finetune-hindi.autotrain-data-skill2go_summ_mbart
AutoTrain Dataset for project: skill2go_summ_mbart
Dataset Description
This dataset has been automatically processed by AutoTrain for project skill2go_summ_mbart.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"feat_Unnamed: 0": 1258,
"text": "<p>\u0414\u0430\u043d\u043d\u044b\u0439 \u043a\u0443\u0440\u0441… See the full description on the dataset page: https://huggingface.co/datasets/PavelDanek/autotrain-data-skill2go_summ_mbart.Evaluation_Lorafacebook-mbart-large-cc25newsdiscusstest2015-fren-100sample-mbart-mariansquad_translated_20k_mbartautotrain-data-mbart_english
Dataset Card for "autotrain-data-mbart_english"
More Information needed
Tagalog_to_Waray_mBARTEvaluation_facebook-mbart-large-cc25mbart50-Finetuned-En-Ml-trans_11Tagalog_to_Ilocano_mBARTtokenized_mBartTagalog_to_Cebuano_mBARTmbart-Large_Tuned_MMLoSo_2025
