AkshitaS/facebook_mlqa_plus
Source Dataset Link: facebook/mlqa Revision: 397ed406c1a7902140303e7faf60fff35b58d285 MLQAMLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance. MLQA consists of over 5K extractive QA instances (12K in English) in SQuAD format in seven languages - English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese. MLQA is highly parallel, with QA instances parallel between 4 different languages on average. MLQA… See the full description on the dataset page: https://huggingface.co/datasets/AkshitaS/facebook_mlqa_plus.
Source Dataset
- Link: facebook/mlqa
- Revision:
397ed406c1a7902140303e7faf60fff35b58d285
MLQA MLQA (MultiLingual Question Answering) is a benchmark dataset for evaluating cross-lingual question answering performance. MLQA consists of over 5K extractive QA instances (12K in English) in SQuAD format in seven languages - English, Arabic, German, Spanish, Hindi, Vietnamese and Simplified Chinese. MLQA is highly parallel, with QA instances parallel between 4 different languages on average.
MLQA Plus MLQA Plus additionally has hin_Latn data generated using indictrans library.
