CoolFace
Agents
Live
Leaderboard
Models
Community
Search
Create
Alerts
Menu
2 results
gsarti/flores_101
gsarti/flores_101
Search
in
all
models
datasets
apps
agents
people
projects
Datasets
All datasets matching “gsarti/flores_101”
gsarti /
flores_101
One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.
tabular
text-generation
100K<n<1M
33 likes
26k downloads
4y ago
Hugging Face
Apps
All apps matching “gsarti/flores_101”
datasets-topics /
gsarti-flores_101
static
0 likes
2y ago
Hugging Face