CoolFace
Modelpublic

facebook/xmod-base

sourceHugging Facemitupdated 3y agoView on Hugging Face
18likes19kdownloads
Model Card

xmod-base

X-MOD is a multilingual masked language model trained on filtered CommonCrawl data containing 81 languages. It was introduced in the paper Lifting the Curse of Multilinguality by Pre-training Modular Transformers (Pfeiffer et al., NAACL 2022) and first released in this repository.

Because it has been pre-trained with language-specific modular components (language adapters), X-MOD differs from previous multilingual models like XLM-R. For fine-tuning, the language adapters in each transformer layer are frozen.

Usage

Tokenizer

This model reuses the tokenizer of XLM-R.

Input Language

Because this model uses language adapters, you need to specify the language of your input so that the correct adapter can be activated:

python
from transformers import XmodModel

model = XmodModel.from_pretrained("facebook/xmod-base")
model.set_default_language("en_XX")

A directory of the language adapters in this model is found at the bottom of this model card.

Fine-tuning

In the experiments in the original paper, the embedding layer and the language adapters are frozen during fine-tuning. A method for doing this is provided in the code:

python
model.freeze_embeddings_and_language_adapters()
# Fine-tune the model ...

Cross-lingual Transfer

After fine-tuning, zero-shot cross-lingual transfer can be tested by activating the language adapter of the target language:

python
model.set_default_language("de_DE")
# Evaluate the model on German examples ...

Bias, Risks, and Limitations

Please refer to the model card of XLM-R, because X-MOD has a similar architecture and has been trained on similar training data.

Citation

BibTeX:

bibtex
@inproceedings{pfeiffer-etal-2022-lifting,
    title = "Lifting the Curse of Multilinguality by Pre-training Modular Transformers",
    author = "Pfeiffer, Jonas  and
      Goyal, Naman  and
      Lin, Xi  and
      Li, Xian  and
      Cross, James  and
      Riedel, Sebastian  and
      Artetxe, Mikel",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.255",
    doi = "10.18653/v1/2022.naacl-main.255",
    pages = "3479--3495"
}

Languages

This model contains the following language adapters:

lang_id (Adapter index)Language codeLanguage
0en_XXEnglish
1id_IDIndonesian
2vi_VNVietnamese
3ru_RURussian
4fa_IRPersian
5sv_SESwedish
6ja_XXJapanese
7fr_XXFrench
8de_DEGerman
9ro_RORomanian
10ko_KRKorean
11hu_HUHungarian
12es_XXSpanish
13fi_FIFinnish
14uk_UAUkrainian
15da_DKDanish
16pt_XXPortuguese
17no_XXNorwegian
18th_THThai
19pl_PLPolish
20bg_BGBulgarian
21nl_XXDutch
22zh_CNChinese (simplified)
23he_ILHebrew
24el_GRGreek
25it_ITItalian
26sk_SKSlovak
27hr_HRCroatian
28tr_TRTurkish
29ar_ARArabic
30cs_CZCzech
31lt_LTLithuanian
32hi_INHindi
33zh_TWChinese (traditional)
34ca_ESCatalan
35ms_MYMalay
36sl_SISlovenian
37lv_LVLatvian
38ta_INTamil
39bn_INBengali
40et_EEEstonian
41az_AZAzerbaijani
42sq_ALAlbanian
43sr_RSSerbian
44kk_KZKazakh
45ka_GEGeorgian
46tl_XXTagalog
47ur_PKUrdu
48is_ISIcelandic
49hy_AMArmenian
50ml_INMalayalam
51mk_MKMacedonian
52be_BYBelarusian
53la_VALatin
54te_INTelugu
55eu_ESBasque
56gl_ESGalician
57mn_MNMongolian
58kn_INKannada
59ne_NPNepali
60sw_KESwahili
61si_LKSinhala
62mr_INMarathi
63af_ZAAfrikaans
64gu_INGujarati
65cy_GBWelsh
66eo_EOEsperanto
67km_KHCentral Khmer
68ky_KGKirghiz
69uz_UZUzbek
70ps_AFPashto
71pa_INPunjabi
72ga_IEIrish
73ha_NGHausa
74am_ETAmharic
75lo_LALao
76ku_TRKurdish
77so_SOSomali
78my_MMBurmese
79or_INOriya
80sa_INSanskrit