PiotrSty/trocr-pl-mixed-v1
031
PiotrSty/trocr-pl-mixed-v1 (experimental)
Fine-tune of PiotrSty/trocr-pl-base on synthetic Polish print + real EHRI typewritten Polish lines (CC-BY 4.0, ehri-pl-lines).
Training
- Base: PiotrSty/trocr-pl-base
- Method: QLoRA, decoder attn q/k/v/out_proj, rank 16, alpha 32
- Train: 2000 synthetic + 349 real EHRI lines (3 docs)
- Val: 38 EHRI lines (held-out doc ZIH3010905)
- Epochs: 5, batch 8, lr 2e-4, T4 x2
- Best checkpoint: checkpoint-735 (val CER 0.3351)
- Document-level split, no line leakage.
Evaluation on frozen held-out sets
Mixed fine-tuning improved BOTH domains (typewriter -28% CER, print -36% CER).
Limitations
- Typewriter WER still ~86%; word-level weak.
- Only 349 real typewritten lines; more data should help.
- Page segmentation on faded typewriter is unreliable; this is a line recognizer.
- Do NOT use as drop-in replacement for trocr-pl-base without your own eval.
Provenance
See run.json, selection.json, bestmetrics.json in this repo. Source: https://github.com/PiotrStyla/OCRengine (commit 7065e6a) EHRI dataset: https://huggingface.co/datasets/PiotrSty/ehri-pl-lines
