CoolFace
Modelpublic

MagnusKolsjo/svensk-fraktur-kraken

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes12downloads
Model Card

Swedish Fraktur — Kraken OCR/HTR model

A Kraken recognition model (.mlmodel) for OCR of older Swedish text printed in Fraktur (blackletter). It produces a diplomatic transcription that keeps historical orthography, the long s (ſ) and period spelling, and works on scanned page images.

  • —File: svensk_fraktur.mlmodel
  • —Type: Kraken recognition model (baseline/HTR pipeline)
  • —DOI: 10.5281/zenodo.20702142
  • —Best validation accuracy: 0.9880 (≈ 1.2 % CER) on a held-out split
  • —Base model: german_print (fine-tuned from it)
  • —Training data: Språkbanken, Svensk fraktur 1626–1816

Usage

Fetch the model directly from Zenodo: kraken get 10.5281/zenodo.20702142

Kraken (CLI):

bash
kraken -i page.jpg out.txt segment -bl ocr -m svensk_fraktur.mlmodel

For multi-page PDFs, render pages to images first (e.g. ~300 DPI) and pass them to Kraken. eScriptorium: import the .mlmodel under Models.

How it was trained

Fine-tuned from german_print on the Språkbanken corpus Svensk fraktur 1626–1816 (199 page images with line-level diplomatic transcriptions). The page images were OCR-bootstrapped and the recognized lines aligned to the ground-truth line text (≈98 % coverage), producing PageXML for ketos train. A low learning rate (-r 0.0001) was decisive — it let the model improve steadily past epoch 0 instead of drifting away from the strong starting point. See `training/` for the scripts and exact commands.

The validation accuracy is measured on a held-out split of the same corpus and is therefore optimistic relative to entirely new documents; on a real volume (1600s Swedish Fraktur) it produced a near-flawless body-text transcription with preserved long-s and period spelling and no systematic substitution errors.

Limitations

  • —Trained on Swedish Fraktur print; not intended for handwriting or modern (antiqua/roman) type.
  • —Ornate/decorated title-page initials are read less reliably than body text.
  • —Output is diplomatic (verbatim): historical spelling, long-s and printed line-break hyphens are preserved. Modernisation/normalisation should happen in a downstream step, not here.

Provenance, rights and attribution

This is a derivative model. Full chain:

LayerResourceByDOILicense
Base modelgerman_print (OCR model for German prints)S. Weil, J. Kamlah, T. Schmidt (2023)10.5281/zenodo.10519596CC0-1.0
Training dataSvensk fraktur 1626–1816Språkbanken Text, University of Gothenburg10.23695/5sme-7437CC-BY-4.0

The base model is CC0 (no attribution legally required; cited as courtesy). The training data is CC-BY-4.0, which requires attribution. When you use or redistribute this model, please keep the following attribution:

Trained on Svensk fraktur 1626–1816, Språkbanken Text, University of Gothenburg (digitisation: Gothenburg University Library; transcription: GREPECT), doi.org/10.23695/5sme-7437, licensed CC-BY-4.0.

This model is released under CC-BY-4.0 (see `LICENSE`).

Citation

See `CITATION.cff`.