CoolFace
Modelpublic

dh-unibe/qwen3vl-german-xix-v2

sourceHugging Faceupdated 4d agoView on Hugging Face
2likes40downloads
Model Card

qwen3vl-german-xix-v2

Handwritten-text-recognition model trained on the `serving-atr-inference` training service. These are the weights of the best validation checkpoint of the run below — not its last epoch.

Evaluation

metricvalue
CER5.33 %
WER18.79 %
samples scored200
characters scored6772
character errors361

Measured on this run's own held-out validation split (page-level and seeded, so no page contributes lines to both sides). It is not a score on a shared benchmark and does not transfer to a different corpus.

The score mixes two kinds of validation, and the difference matters. dh-unibe/image-text_kurrent-xix held whole projects out of training, so those lines test unseen hands. dh-unibe/image-text_zh-regierungsratsprotokolle, dh-unibe/image-text_parlamentsdienste-protokolle, dh-unibe/image-text_nr-sr-vereinigte-bundesversammlung-xix contributed a seeded partition of their own training projects instead — unseen pages in a hand the model trained on, which is the easier test. The figure above is one CER over both, so read it as mostly in-domain, not as a held-out-hands benchmark. Scoring the held-out projects on their own would give the stricter number.

Input granularity: give it one line, not a page

Trained on line crops (granularity: line, 262 144 pixels), and it reads that way. Measured 2026-09-22 with `scripts/eval_granularity.py` on 15 validation pages of this run, against their ground truth:

inputnCERlength ratio
line crops6530.0521.00
paragraphs (TextRegions)360.950.05
whole pages150.980.02

A page holding 376-5674 characters comes back as 17-60. What comes back is a fluent German line that is often not on the page at all — "Hochzeitlich in der Stadt" for a page beginning "s¬ Wyß, als dem Herrn Bezirks¬ statthalter Steiner". That is harder to notice than an obvious failure, so: segment the page first (kraken, or any line segmenter) and send one crop per call. No pixel budget fixes the page shape; the model learned to stop after one line.

Details: thodel/serving-atr-inference#165.

Training data

  • —`dh-unibe/image-text_zh-regierungsratsprotokolle`
  • —Training projects: —
  • —Evaluation: a seeded page-level split of the training projects (partition=0.9, seed=42)
  • —Page cap: 20000
  • —`dh-unibe/image-text_parlamentsdienste-protokolle`
  • —Training projects: 1849_01, 1849_02, 1849_03, 1849_04, 1849_05, 1849_06, 1849_07, 1849_08, 637157213248534270_verkleinert
  • —Evaluation: a seeded page-level split of the training projects (partition=0.9, seed=42)
  • —Page cap: 138
  • —`dh-unibe/image-text_nr-sr-vereinigte-bundesversammlung-xix`
  • —Training projects: Nationalrat_01__Sitzung_–_06_11_1848, Nationalrat_02__Sitzung_–_07_11_1848, Nationalrat_03__Sitzung_–_08_11_1848, Nationalrat_04__Sitzung_–_09_11_1848, Nationalrat_05__Sitzung_–_10_11_1848, Nationalrat_06__Sitzung_–_13_11_1848, Nationalrat_07__Sitzung_–_14_11_1848, Nationalrat_08__Sitzung_–_15_11_1848, Nationalrat_09__Sitzung_–_18_11_1848, Nationalrat_10__Sitzung_–_21_11_1848, Nationalrat_11__Sitzung_–_22_11_1848, Nationalrat_12__Sitzung_–_23_11_1848, Nationalrat_13__Sitzung_–_24_11_1848, Nationalrat_14__Sitzung_–_25_11_1848, Nationalrat_15__Sitzung_–_27_11_1848, Nationalrat_16__Sitzung_–_28_11_1848, Nationalrat_17__Sitzung_–_29_11_1848, Ständerat_01__Sitzung_–_06_11_1848, Ständerat_02__Sitzung_–_08_11_1848, Ständerat_03__Sitzung_–_09_11_1848, Ständerat_04__Sitzung_–_11_11_1848, Ständerat_05__Sitzung_–_14_11_1848, Ständerat_06__Sitzung_–_15_11_1848, Ständerat_07__Sitzung_–_21_11_1848, Ständerat_08__Sitzung_–_23_11_1848, Ständerat_09__Sitzung_–_24_11_1848, Ständerat_10__Sitzung_–_25_11_1848, Ständerat_11__Sitzung_–_27_11_1848, Ständerat_12__Sitzung_–_28_11_1848, Ständerat_13__Sitzung_–_28_11_1848, VBV_01__Sitzung_–_16_11_1848, VBV_02__Sitzung_–_17_11_1848, VBV_03__Sitzung_–_20_11_1848, VBV_04__Sitzung_–_24_11_1848, VBV_05__Sitzung_–_29_11_1848
  • —Evaluation: a seeded page-level split of the training projects (partition=0.9, seed=42)
  • —Page cap: 52
  • —`dh-unibe/image-text_kurrent-xix`
  • —Training projects: TRAIN_CITlab_Bassermann_0_4, TRAIN_CITlab_Bassermann_Manuscripts, TRAIN_CITlab_Bassermann_Manuscripts_0_2, TRAIN_CITlab_Binder_Kochbuch_2, TRAIN_CITlab_Escher_M1, TRAIN_CITlab_Gusbeth, TRAIN_CITlab_Handschriftliche_Archivquellen, TRAIN_CITlab_Konzilsprotkolle_B_Schwartz_5_2018, TRAIN_CITlab_MargareteSick, TRAIN_CITlab_MargareteSick20180731, TRAIN_CITlab_MargareteSick20180731a, TRAIN_CITlab_MÜLLER, TRAIN_CITlab_Protokoll_Hoftheater_1806, TRAIN_CITlab_Rehlen_1834_a, TRAIN_CITlab_RvE_Barlaam_HS_D_, TRAIN_CITlab_Steiner, TRAIN_CITlab_Suppes, TRAIN_CITlab_Suppes_3500, TRAIN_CITlab_Tagebuch_Arnold_v1, TRAIN_CITlab_umkc_Roland_M1, TRAIN_CITlab_umkc_Roland_M2, hufeland_privatbesitz_1829, nn_msgermqu2124_1827, nn_msgermqu2345_1827, parthey
  • —Evaluation: held-out projects TEST_CITlab_Bassermann_0_4, TEST_CITlab_Bassermann_Manuscripts, TEST_CITlab_Bassermann_Manuscripts_0_2, TEST_CITlab_Binder_Kochbuch_2, TEST_CITlab_Escher_M1, TEST_CITlab_Gusbeth, TEST_CITlab_Handschriftliche_Archivquellen, TEST_CITlab_Konzilsprotkolle_B_Schwartz_5_2018, TEST_CITlab_MargareteSick, TEST_CITlab_MargareteSick20180731, TEST_CITlab_MargareteSick20180731a, TEST_CITlab_MÜLLER, TEST_CITlab_Protokoll_Hoftheater_1806, TEST_CITlab_Rehlen_1834_a, TEST_CITlab_RvE_Barlaam_HS_D_, TEST_CITlab_Steiner, TEST_CITlab_Suppes, TEST_CITlab_Suppes_3500, TEST_CITlab_Tagebuch_Arnold_v1, TEST_CITlab_umkc_Roland_M1, TEST_CITlab_umkc_Roland_M2
  • —Page cap: 19808

Materialized from that selection: 26,186 pages, 964,472 transcribed lines, 966,748 training samples.

Trained with the instruction: Transcribe the handwritten text in this image exactly as written. — serving it with different wording is a silent distribution shift.

Hyperparameters

yaml
granularity: line
prompt: Transcribe the handwritten text in this image exactly as written.
load_in_4bit: false
lora_r: 64
lora_alpha: 128
lora_dropout: 0.05
target_modules:
- q_proj
- k_proj
- v_proj
- o_proj
- gate_proj
- up_proj
- down_proj
modules_to_save: []
epochs: 1
max_epochs: null
patience: 2
min_delta: 0.0001
batch_size: 16
accumulate_grad_batches: 1
lrate: 0.0002
lr_scheduler: cosine
warmup_ratio: 0.05
weight_decay: 0.0
max_grad_norm: 1.0
optim: paged_adamw_8bit
gradient_checkpointing: true
save_steps: 200
max_pixels: 262144
max_seq_len: 1024
min_train_chars: 0
eval_samples: 200
max_new_tokens: null
seed: 42
workers: 8
device: cuda:0
wandb_run: null

Provenance

enginevllm
base model`Qwen/Qwen3-VL-4B-Instruct`
training job20260916T090417Z-qwen3vl-german-xix-v2
codenot recorded (trained before #147)
trained2026-09-17T16:34:54.094758+00:00
weightsadapter_config.json, adapter_model.safetensors, added_tokens.json, chat_template.jinja, merges.txt, preprocessor_config.json, special_tokens_map.json, tokenizer.json, tokenizer_config.json, training_summary.json, video_preprocessor_config.json, vocab.json

metadata.json in this repo is the record the trainer wrote, verbatim: the full request, the parsed metrics and the job id.

Using it

This is a LoRA adapter, not a full model — it needs its base:

python
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor

base = AutoModelForImageTextToText.from_pretrained('Qwen/Qwen3-VL-4B-Instruct')
model = PeftModel.from_pretrained(base, 'dh-unibe/qwen3vl-german-xix-v2')
processor = AutoProcessor.from_pretrained('dh-unibe/qwen3vl-german-xix-v2', trust_remote_code=True)

vLLM 0.11 will not serve it as an adapter (it refuses LoRA on the vision tower), so serving means merging it into the base first — scripts/merge_loras.py in serving-atr-inference does that.

Notes

HELD-OUT RESULT. On the published benchmark 'Handwritten Text Recognition Test Set: Minutes of the Swiss Federal Council (1848-1903)' (Hodel & Schoch 2021, Zenodo, https://doi.org/10.5281/zenodo.4746342; private HF mirror dh-unibe/image-textfederal-minutes-testset), all 2,751 lines, no document shared with training: CER 0.0765, WER 0.2458, lengthratio 1.0019.

WHAT THIS REPLACES. qwen3vl-german-xix-v1 scored CER 0.2551 on the same lines and wrote little more than the first word on 507 of them (18.4%). v2 collapses on 0. The only difference is the corpus: v1 was built before the PageXML fix 33f55fc (#125), which had truncated 6.35x the characters of nr-sr-vereinigte-bundesversammlung-xix and 4.54x of parlamentsdienste-protokolle to their first word. Same four repositories, same seed, same page-level split.

THE CER IN THE TABLE ABOVE (0.0533) IS THE VALIDATION SPLIT, not the benchmark, and it is NOT comparable with v1's 0.0100: v1 was scored on the first 200 lines of val.jsonl (five in-domain pages), v2 on the stratified draw introduced in #120. Quote 0.0765 as this model's accuracy on unseen hands, and 0.0533 only as its score on its own corpus.

The remaining errors are 5,499 substitutions against 1,538 missing characters - misreadings rather than lost text.