CoolFace
Modelpublic

clarin-pl/combo-seg-xlm-roberta-base-norwegian-nynorsk-ud2.17

sourceHugging Facecc-by-sa-4.0updated 2mo agoView on Hugging Face
0likes18downloads
Model Card

COMBO-SEG Model for Norwegian

Model Description

This is a Norwegian-language character-level segmentation model based on COMBO-SEG, an open-source text segmentation system. It performs:

  • —sentence segmentation
  • —tokenisation (including multi-word token detection)

The Norwegian model uses `FacebookAI/xlm-roberta-base` as its base encoder and is trained on UD_Norwegian-Nynorsk (UD v2.17).

Evaluation

MetricTokensWordsSentences
F199.9699.9697.74

Usage

Install the library from PyPI:

bash
pip install combo-seg
python
from combo_seg import ComboSeg

# Load a pre-trained model
nlp = ComboSeg("Norwegian")

# Segment raw text — returns Document with hierarchy: Document -> Turn -> Sentence -> Token
doc = nlp("Den raske brune reven hopper over den late hunden.")

# Inspect results
for turn in doc.turns:
    for sentence in turn.sentences:
        print(f"Sentence: {sentence.text}")
        for token in sentence.tokens:
            if token.is_multi_word:
                print(f"  MWT: {token.text} -> {token.subwords}")
            else:
                print(f"  Token: {token.text}")

Or load directly from HuggingFace:

python
from combo_seg import ComboSeg

nlp = ComboSeg.from_pretrained("clarin-pl/combo-seg-xlm-roberta-base-norwegian-nynorsk-ud2.17")
doc = nlp("Den raske brune reven hopper over den late hunden.")

License

The training data license: cc-by-sa-4.0 is derived from the Universal Dependencies treebank. For the full license terms of each treebank, please refer to the corresponding LICENSE.txt file in the treebank repository:

Citation

If you use this model, please cite:

Ulewicz, M., & Wróblewska, A. (2026). COMBO-SEG Models Trained on UD v2.17. https://doi.org/10.5281/zenodo.19651441

bibtex
@software{combo_seg_2026,
  author    = {Ulewicz, Michał and Wróblewska, Alina},
  title     = {{COMBO-SEG} Models Trained on {UD} v2.17},
  year      = {2026},
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.19651441},
  url       = {https://doi.org/10.5281/zenodo.19651441}
}

Resources