CoolFace
Datasetpublic

Helsinki-NLP/fineweb-edu-translated

Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/fineweb-edu-translated.

sourceHugging Faceodc-byupdated 5mo agoView on Hugging Face
16likes186kdownloads
Dataset Card

Helsinki-NLP/fineweb-edu-translated

fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages.

In the v1.1 release, additional translations are added for Czech (ces), Ukrainian (ukr) and Finnish (fin). For Czech and Ukrainian, this release doubles the data and for Finnish, we include translations for the entire fineweb-edu data set with its 350B token release.

More information about how the data has been produced can be found on https://github.com/Helsinki-NLP/translate-fineweb. The corpus is also available with aligned sentences in synOPUS: https://opus.nlpl.eu/synthetic/transweb-edu.php

Supported Languages

LangidLanguage
bosBosnian
bulBulgarian
catCatalan
cesCzech
danDanish
deuGerman
ellModern Greek
engEnglish
estEstonian
eusBasque
finFinnish
fraFrench
gleIrish
glgGalician
hrvCroatian
hunHungarian
islIcelandic
itaItalian
katGeorgian
lavLatvian
litLithuanian
mkdMacedonian
mltMaltese
nldDutch
nnoNorwegian Nynorsk
nobNorwegian Bokmål
polPolish
porPortuguese
ronRomanian
slkSlovak
slvSlovenian
spaSpanish
sqiAlbanian
srp_CyrlSerbian (cyrillic script)
sweSwedish
turTurkish
ukrUkrainian

Citation Information

Please acknowledge the source when using the data and, please, cite the following article if you use any part of this corpus in your own work:

bibtex
@article{tiedemann2023democratizing,
  title={Democratizing neural machine translation with {OPUS-MT}},
  author={Tiedemann, J{\"o}rg and Aulamo, Mikko and Bakshandaeva, Daria and Boggia, Michele and Gr{\"o}nroos, Stig-Arne and Nieminen, Tommi and Raganato, Alessandro and Scherrer, Yves and Vazquez, Raul and Virpioja, Sami},
  journal={Language Resources and Evaluation},
  number={58},
  pages={713--755},
  year={2023},
  publisher={Springer Nature},
  issn={1574-0218},
  doi={10.1007/s10579-023-09704-w}
}

Translation Models

The following translation models have been used for creating the data:

Langidtranslation model
bosTatoeba-MT-models/eng-hbs/opus+bt-2021-04-20download
bulTatoeba-MT-models/eng-bul/opusTCv20210807+bt_transformer-big_2022-02-25download
catTatoeba-MT-models/deu+eng+fra+por+spa-roa/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
cesTatoeba-MT-models/eng-ces+slk/opusTCv20210807+bt_transformer-big_2022-03-13download
danTatoeba-MT-models/eng-gmq/opusTCv20210807+bt_transformer-big_2022-03-17download
deuTatoeba-MT-models/eng-deu/opusTCv20210807+bt-2021-12-08download
ellTatoeba-MT-models/eng-ell/opusTCv20210807+bt_transformer-big_2022-03-13download
estTatoeba-MT-models/deu+eng+fra+por+spa-urj/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
eusHPLT-MT-models/en-eu/translate-en-eu-v1.0-hplt_opusdownload
finTatoeba-MT-models/eng-fin/opusTCv20210807+bt_transformer-big_2022-03-09download
fraTatoeba-MT-models/eng-fra/opusTCv20210807+bt_transformer-big_2022-03-09download
gleHPLT-MT-models/en-ga/translate-en-ga-v1.0-hplt_opusdownload
glgTatoeba-MT-models/deu+eng+fra+por+spa-itc/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
hrvTatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
hunTatoeba-MT-models/eng-hun/opusTCv20210807+bt_transformer-big_2022-02-25download
islHPLT-MT-models/en-is/translate-en-is-v1.0-hplt_opusdownload
itaTatoeba-MT-models/deu+eng+fra+por+spa-itc/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
katTatoeba-MT-models/deu+eng+fra+por+spa-cau/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
lavTatoeba-MT-models/eng-lav/opusTCv20210807+bt_transformer-big_2022-03-13download
litTatoeba-MT-models/deu+eng+fra+por+spa-bat/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
mkdTatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
mltHPLT-MT-models/en-mt/translate-en-mt-v1.0-hplt_opusdownload
nldTatoeba-MT-models/deu+eng+fra+por+spa-gmw/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
nnoTatoeba-MT-models/deu+eng+fra+por+spa-gmq/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
nobTatoeba-MT-models/deu+eng+fra+por+spa-gmq/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
polTatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
porTatoeba-MT-models/gem-fra+ita+por+spa/opusTCv20230926max50+bt+jhubc_transformer-big_2024-08-17download
ronTatoeba-MT-models/deu+eng+fra+por+spa-itc/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
slkTatoeba-MT-models/eng-ces+slk/opusTCv20210807+bt_transformer-big_2022-03-13download
slvTatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
spaTatoeba-MT-models/eng-spa/opusTCv20210807+bt_transformer-big_2022-03-13download
sqiHPLT-MT-models/en-sq/translate-en-sq-v1.0-hplt_opusdownload
srp_CyrlTatoeba-MT-models/deu+eng+fra+por+spa-sla/opusTCv20230926max50+bt+jhubc_transformer-big_2024-05-30download
sweTatoeba-MT-models/eng-gmq/opusTCv20210807+bt_transformer-big_2022-03-17download
turTatoeba-MT-models/eng-tur/opusTCv20210807+bt_transformer-big_2022-02-25download
ukrTatoeba-MT-models/eng-zle/opusTCv20210807+bt_transformer-big_2022-03-13download

Acknowledgements

None of this would be possible without the enormous work done by Common Crawl providing the essential data that most open datasets for language modeling are based on. Furthermore, we are also grateful for the data preparation work done by Hugging Face and the community on top of the crawled data from Common Crawl published under the label of fineweb-edu. Important for this work is also the availability of parallel data through OPUS and the public translation models based on that data. Furthermore, this project was supported by the European Union's Horizon Europe research and innovation programme through the HPLT project under grant agreement No 101070350. Finally, we also want to acknowledge the computational resource made available from the Finnish national allocation for the LUMI supercomputer (https://www.lumi-supercomputer.eu) through the extreme scale project MaMuLaM: Massively Multilingual Language Models. None of the translated data would exist without this infrastructure and the compute.