CoolFace
Modelpublic

hmteams/teams-base-historic-multilingual-discriminator

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes70downloads
Model Card

hmTEAMS

![🤗](https://github.com/stefan-it/hmTEAMS)

Historic Multilingual and Monolingual TEAMS Models. The following languages are covered:

  • —English (British Library Corpus - Books)
  • —German (Europeana Newspaper)
  • —French (Europeana Newspaper)
  • —Finnish (Europeana Newspaper, Digilib)
  • —Swedish (Europeana Newspaper, Digilib)
  • —Dutch (Delpher Corpus)
  • —Norwegian (NCC Corpus)

Architecture

We pretrain a "Training ELECTRA Augmented with Multi-word Selection" (TEAMS) model:

hmTEAMS Overview

Results

We perform experiments on various historic NER datasets, such as HIPE-2022 or ICDAR Europeana. All details incl. hyper-parameters can be found here.

Small Benchmark

We test our pretrained language models on various datasets from HIPE-2020, HIPE-2022 and Europeana. The following table shows an overview of used datasets.

LanguageDatasetAdditional Dataset
EnglishAjMC-
GermanAjMC-
FrenchAjMCICDAR-Europeana
FinnishNewsEye-
SwedishNewsEye-
DutchICDAR-Europeana-

Results

ModelEnglish AjMCGerman AjMCFrench AjMCFinnish NewsEyeSwedish NewsEyeDutch ICDARFrench ICDARAvg.
hmBERT (32k) Schweter et al.85.36 ± 0.9489.08 ± 0.0985.10 ± 0.6077.28 ± 0.3782.85 ± 0.8382.11 ± 0.6177.21 ± 0.1682.71
hmTEAMS (Ours)86.41 ± 0.3688.64 ± 0.4285.41 ± 0.6779.27 ± 1.8882.78 ± 0.6088.21 ± 0.3978.03 ± 0.3984.11

Release

Our pretrained hmTEAMS model can be obtained from the Hugging Face Model Hub:

Acknowledgements

We thank Luisa März, Katharina Schmid and Erion Çano for their fruitful discussions about Historic Language Models.

Research supported with Cloud TPUs from Google's TPU Research Cloud (TRC). Many Thanks for providing access to the TPUs ❤️