datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_plus
Dataset Card for FLORES+
FLORES+ is an evaluation benchmark dataset for multilingual machine translation.
Dataset Details
Dataset Description
FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.flores-plusplus-blocks
FLORES++ blocks
Blocks of k = 1..5 consecutive segments of one FLORES+ article (dev + devtest),
English source, with the correct translation (the reference) and its
distractors: copies of the reference in which every segment carries one
perturbation from the three xSIM++ categories (causality, entity, number;
Chen et al., 2023), in every possible combination (up to 243 at k = 5).
The error share is one per segment at every k, so if a model finds the reference less
often on… See the full description on the dataset page: https://huggingface.co/datasets/AdleBenSalem/flores-plusplus-blocks.
