Helsinki-NLP/opus-mt-tc-bible-big-mul-mul
opus-mt-tc-bible-big-mul-mul
Table of Contents
- Model Details
- Uses
- Risks, Limitations and Biases
- How to Get Started With the Model
- Training
- Evaluation
- Citation Information
- Acknowledgements
Model Details
Neural machine translation model for translating from Multiple languages (mul) to Multiple languages (mul). Note that many of the listed languages will not be well supported by the model as the training data is very limited for the majority of the languages. Translation performance varies a lot and for a large number of language pairs it will not work at all.
This model is part of the OPUS-MT project, an effort to make neural machine translation models widely available and accessible for many languages in the world. All models are originally trained using the amazing framework of Marian NMT, an efficient NMT implementation written in pure C++. The models have been converted to pyTorch using the transformers library by huggingface. Training data is taken from OPUS and training pipelines use the procedures of OPUS-MT-train. Model Description:
- Developed by: Language Technology Research Group at the University of Helsinki
- Model Type: Translation (transformer-big)
- Release: 2024-08-17
- License: Apache-2.0
- Language(s):
- Source Language(s): aar abk ace ach acm ady afb afh afr aii ajp aka akl aln alt amh ami amu ang anp aoz apc ara arc arg arq arz asm ast atj ava avk awa ayl aze azz bak bal bam ban bar bas bcl bel bem ben bho bik bis bod bom bos bpy bre brx bua bug bul bvy byn bzt cak cat cay cbk ceb ces cha che chg chm chq chr chu chv chy cjk cjp cjy ckb cmn cnh cni cnr cop cor cos cre crh crk crs csb cym dag dan deu dik din diq div dje djk dng dop drt dsb dtp dty dws dyu dzo efi egl ell emx eng enm epo est eus evn ewe ext fao fas fij fil fin fkv fon fra frm fro frp frr fry fuc ful fur gag gbm gcf gil gla gle glg glk glv gor gos got grc grn gsw guc guj guw hat hau haw hbo hbs heb her hif hil hin hmn hne hnj hoc hrv hrx hsb hsn hun hus hye hyw iba ibo ido igs iii ike iku ile ilo ina ind inh ipk isl ita ixl izh jaa jak jam jav jbo jdt jpa jpn kaa kab kac kal kam kan kas kat kau kaz kbd kbp kea kek kha khm kik kin kir kiu kjh kmb kmr knc koi kok kom kon kpv krc krl ksh kua kum kur kxi laa lad lah lao lat lav lbe ldn lez lfn lij lim lin lit liv lkt lld lmo lou lrc ltz lua lug luo lus lut luy lzz mad mag mah mai mal mam mar max mdf meh mfa mfe mgm mic mix mkd mlg mlt mnc mni mnr mnw moh mol mon mos mri mrj msa mvv mwl mww mya myv mzn nap nau nav nbl nch nde nds nep new ngt ngu nhg nhn nia niu nld nlv nnb nno nob nog non nov npi nqo nso nst nus nya oar oci ofs oji ood ori orm orv osp oss ota otk pag pai pal pam pan pap pau pcd pck pcm pdc pes pfl phn pih pli plt pms pmy pnt pol por pot ppk ppl prg prs pus quc qxq qya rap rhg rif rmy roh rom ron rue run rup rus sag sah san sat scn sco sdh ses sgs shi shn shs shy sin sjn skr slk slv sma sme sml smn smo sna snd som sot spa sqi srd srn srp ssw stq sun swa swc swe swg swh syc syl syr szl tah tam taq tat tcy tel tet tgk tgl tha thv tig tir tkl tlh tly tmh tmr tmw toi ton tpi tpw trs trv tsn tso tts tuk tum tur tvl twi tyj tyv tzl tzm udm uig ukr umb urd usp uzb vec ven vep vie vls vol vot vro wae wal war wln wol wuu xal xcl xho xmf yid yor yua yue zam zap zea zgh zha zlm zsm zul zza
- Target Language(s): aar abk ace ach acm ady afb afhLatn afr aiiSyrc ajp aka aklLatn aln alt amh ami amiLatn amuLatn angLatn anp aoz apc ara arc arg arq arz asm ast atj ava avkLatn awa ayl azeCyrl azeLatn azz azzLatn bak bal balLatn bamLatn ban bar bas bcl bel bem ben bho bik bis bod bomLatn bosCyrl bosLatn bpy bre brx bua bug bul bvyLatn byn bztLatn cak cakLatn cat cay cbkLatn ceb ces cha che chgArab chgLatn chm chqLatn chr chu chv chy cjk cjkLatn cjpLatn cjyHans cjyHant ckb cmn cmnHans cmnHant cnh cnhLatn cniLatn cnr cnrLatn cop copCopt cor cos cre creLatn crh crk crs csb csbLatn cym dagLatn dan deu dik din diq div dje djk djkLatn dng dopLatn drtLatn dsb dtp dty dwsLatn dyu dzo efi egl ell emxLatn eng enmLatn epo est eus evn ewe ext fao fas fij fil fin fkvLatn fon fra frmLatn froLatn frp frr fry fuc ful fur gag gbm gcf gcfLatn gil gla gle glg glk glv gor gos got gotGoth grc grcGrek grn gsw guc guj guw guwLatn hat hauLatn haw hboHebr hbs hbsCyrl hbsLatn heb her hifLatn hil hin hinLatn hmn hne hnj hoc hocWara hrv hrxLatn hsb hsn hun hus husLatn hye hyw hywArmn hywLatn iba ibo idoLatn igsLatn iii ikeLatn ikuLatn ile ileLatn ilo inaLatn ind inh inhLatn ipk isl ita ixlLatn izh jaa jaaBopo jaaHira jaaKana jaaYiii jakLatn jam jav javJava jbo jboCyrl jboLatn jdtCyrl jpaHebr jpn kaa kab kac kal kam kan kasArab kasDeva kat kau kaz kazCyrl kbd kbp kbpCans kbpEthi kbpGeor kbpGrek kbpHang kbpLatn kbpMlym kbpYiii kea kek kekLatn kha khm kik kin kirCyrl kiu kjh kmb kmr knc koi kok kom kon kpv krc krl ksh kua kum kurArab kurCyrl kurLatn kxiLatn laaLatn lad ladLatn lah lao lat latLatn lav lbe ldnLatn lez lfnCyrl lfnLatn lij lim lin lit livLatn lkt lldLatn lmo louLatn lrc ltz lua lug luo lus lutLatn luy lzzGeor lzzLatn mad mag mah mai mal mam mamLatn mar maxLatn mdf mehLatn mfa mfe mgmLatn mic mix mixLatn mkd mlg mlt mncMong mni mnrLatn mnw moh mol mon mos mri mrj msaArab msaLatn mvvLatn mwl mww mya myv mzn nap nau nav nbl nch nde nds nep new ngtLatn ngu nguLatn nhgLatn nhnLatn nia niu nld nlvLatn nnbLatn nno nob nog non novLatn npi nqo nso nstLatn nus nya oarHebr oarSyrc oci ofsLatn ojiLatn oodLatn ori orm orvCyrl ospLatn oss otaArab otaLatn otaRohg otaSyrc otaThaa otaYezi otk otkOrkh pag paiLatn pal pam pan panGuru pap pau pcd pckLatn pcm pdc pes pfl phnPhnx pih pihLatn pli plt pms pmyLatn pntGrek pol por potLatn ppkLatn pplLatn prgLatn prs pus quc qxqArab qxqLatn qya qyaLatn rap rhgLatn rifLatn rmy roh rom romCyrl ron rue run rup rus sag sah san sanDeva sat satLatn scn sco sdh ses sgs shiLatn shn shsLatn shyLatn sin sjnLatn skr slk slv sma sme smlLatn smn smo sna sndArab som sot spa sqi srd srn srpCyrl ssw stq sun swa swc swe swg swh sycSyrc sylSylo syr szl tah tam taq tat tcy tel tet tgkCyrl tgkLatn tgl tglLatn tglTglg tha thv tig tir tkl tlh tlhLatn tlyLatn tmh tmrHebr tmwLatn toi toiLatn ton tpi tpwLatn trs trsLatn trv tsn tso tts tuk tukCyrl tukLatn tum tur tvl twi tyjLatn tyv tzl tzlLatn tzmLatn tzmTfng udm uig uigArab uigCyrl uigLatn ukr umb urd uspLatn uzbCyrl uzbLatn vec ven vep vie vls volLatn vot votLatn vro wae wal war wln wol wuu xal xclArmn xclLatn xho xmf yid yor yua yueHans yueHant zam zap zea zgh zha zlmArab zlmLatn zsmArab zsm_Latn zul zza
- Original Model: opusTCv20230926+bt+jhubc_transformer-big_2024-08-17.zip
- Resources for more information:
- OPUS-MT dashboard
- OPUS-MT-train GitHub Repo
- More information about MarianNMT models in the transformers library
- Tatoeba Translation Challenge
- HPLT bilingual data v1 (as part of the Tatoeba Translation Challenge dataset)
- A massively parallel Bible corpus
This is a multilingual translation model with multiple target languages. A sentence initial language token is required in the form of >>id<< (id = valid target language ID), e.g. >>aar<<
Uses
This model can be used for translation and text-to-text generation.
Risks, Limitations and Biases
CONTENT WARNING: Readers should be aware that the model is trained on various public data sets that may contain content that is disturbing, offensive, and can propagate historical and current stereotypes.
Significant research has explored bias and fairness issues with language models (see, e.g., Sheng et al. (2021) and Bender et al. (2021)).
Also note that many of the listed languages will not be well supported by the model as the training data is very limited for the majority of the languages. Translation performance varies a lot and for a large number of language pairs it will not work at all.
How to Get Started With the Model
A short example code:
from transformers import MarianMTModel, MarianTokenizer
src_text = [
">>rus<< You'd better not speak to Tom about that.",
">>ceb<< How are you?"
]
model_name = "pytorch-models/opus-mt-tc-bible-big-mul-mul"
tokenizer = MarianTokenizer.from_pretrained(model_name)
model = MarianMTModel.from_pretrained(model_name)
translated = model.generate(**tokenizer(src_text, return_tensors="pt", padding=True))
for t in translated:
print( tokenizer.decode(t, skip_special_tokens=True) )
# expected output:
# Лучше бы не поговорить с Томом об этом.
# Sa unsang paagi ikaw?You can also use OPUS-MT models with the transformers pipelines, for example:
from transformers import pipeline
pipe = pipeline("translation", model="Helsinki-NLP/opus-mt-tc-bible-big-mul-mul")
print(pipe(">>rus<< You'd better not speak to Tom about that."))
# expected output: Лучше бы не поговорить с Томом об этом.Training
- Data: opusTCv20230926+bt+jhubc (source)
- Pre-processing: SentencePiece (spm64k,spm64k)
- Model Type: transformer-big
- Original MarianNMT Model: opusTCv20230926+bt+jhubc_transformer-big_2024-08-17.zip
- Training Scripts: GitHub Repo
Evaluation
- Model scores at the OPUS-MT dashboard
- test set translations: opusTCv20230926+bt+jhubc_transformer-big_2024-08-17.test.txt
- test set scores: opusTCv20230926+bt+jhubc_transformer-big_2024-08-17.eval.txt
- benchmark results: benchmark_results.txt
- benchmark output: benchmark_translations.zip
Citation Information
- Publications: Democratizing neural machine translation with OPUS-MT and OPUS-MT – Building open translation services for the World and The Tatoeba Translation Challenge – Realistic Data Sets for Low Resource and Multilingual MT (Please, cite if you use this model.)
@article{tiedemann2023democratizing,
title={Democratizing neural machine translation with {OPUS-MT}},
author={Tiedemann, J{\"o}rg and Aulamo, Mikko and Bakshandaeva, Daria and Boggia, Michele and Gr{\"o}nroos, Stig-Arne and Nieminen, Tommi and Raganato, Alessandro and Scherrer, Yves and Vazquez, Raul and Virpioja, Sami},
journal={Language Resources and Evaluation},
number={58},
pages={713--755},
year={2023},
publisher={Springer Nature},
issn={1574-0218},
doi={10.1007/s10579-023-09704-w}
}
@inproceedings{tiedemann-thottingal-2020-opus,
title = "{OPUS}-{MT} {--} Building open translation services for the World",
author = {Tiedemann, J{\"o}rg and Thottingal, Santhosh},
booktitle = "Proceedings of the 22nd Annual Conference of the European Association for Machine Translation",
month = nov,
year = "2020",
address = "Lisboa, Portugal",
publisher = "European Association for Machine Translation",
url = "https://aclanthology.org/2020.eamt-1.61",
pages = "479--480",
}
@inproceedings{tiedemann-2020-tatoeba,
title = "The Tatoeba Translation Challenge {--} Realistic Data Sets for Low Resource and Multilingual {MT}",
author = {Tiedemann, J{\"o}rg},
booktitle = "Proceedings of the Fifth Conference on Machine Translation",
month = nov,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2020.wmt-1.139",
pages = "1174--1182",
}Acknowledgements
The work is supported by the HPLT project, funded by the European Union’s Horizon Europe research and innovation programme under grant agreement No 101070350. We are also grateful for the generous computational resources and IT infrastructure provided by CSC -- IT Center for Science, Finland, and the EuroHPC supercomputer LUMI.
Model conversion info
- transformers version: 4.45.1
- OPUS-MT git hash: 0882077
- port time: Wed Oct 9 19:20:34 EEST 2024
- port machine: LM0-400-22516.local
