CoolFace
Datasetpublic

learningmachineaz/translate_enaz_10m

Machine translation EN-AZ dataset based on Google Translate and National Library of Azerbaijan.

sourceHugging Faceopenrailupdated 3y agoView on Hugging Face
3likes12downloads
Dataset Card

Description

Dataset used to train our mT5 based model for machine translation, extracted from various text sources of National Library of Azerbaijan: mT5-translation-enaz \ It has only clean texts. Wiki articles wasn't used as they contain a lot of irrelevant data.

Key pointInfo
Rows~10mil. EN-AZ sentence pairs
Size975M (zipped) / 2.8G (unzipped)
FormatTSV (tab separated pairs)
EnglishGoogle Translate
AzerbaijaniOriginal cleaned text

Author

Collected and prepared by Renat Kalimulin