dh-unibe/qwen3vl-german-xix-v2
qwen3vl-german-xix-v2
Handwritten-text-recognition model trained on the `serving-atr-inference` training service. These are the weights of the best validation checkpoint of the run below — not its last epoch.
Evaluation
Measured on this run's own held-out validation split (page-level and seeded, so no page contributes lines to both sides). It is not a score on a shared benchmark and does not transfer to a different corpus.
The score mixes two kinds of validation, and the difference matters. dh-unibe/image-text_kurrent-xix held whole projects out of training, so those lines test unseen hands. dh-unibe/image-text_zh-regierungsratsprotokolle, dh-unibe/image-text_parlamentsdienste-protokolle, dh-unibe/image-text_nr-sr-vereinigte-bundesversammlung-xix contributed a seeded partition of their own training projects instead — unseen pages in a hand the model trained on, which is the easier test. The figure above is one CER over both, so read it as mostly in-domain, not as a held-out-hands benchmark. Scoring the held-out projects on their own would give the stricter number.
Input granularity: give it one line, not a page
Trained on line crops (granularity: line, 262 144 pixels), and it reads that way. Measured 2026-09-22 with `scripts/eval_granularity.py` on 15 validation pages of this run, against their ground truth:
A page holding 376-5674 characters comes back as 17-60. What comes back is a fluent German line that is often not on the page at all — "Hochzeitlich in der Stadt" for a page beginning "s¬ Wyß, als dem Herrn Bezirks¬ statthalter Steiner". That is harder to notice than an obvious failure, so: segment the page first (kraken, or any line segmenter) and send one crop per call. No pixel budget fixes the page shape; the model learned to stop after one line.
Details: thodel/serving-atr-inference#165.
Training data
- `dh-unibe/image-text_zh-regierungsratsprotokolle`
- Training projects: —
- Evaluation: a seeded page-level split of the training projects (
partition=0.9,seed=42) - Page cap: 20000
- `dh-unibe/image-text_parlamentsdienste-protokolle`
- Training projects:
1849_01,1849_02,1849_03,1849_04,1849_05,1849_06,1849_07,1849_08,637157213248534270_verkleinert - Evaluation: a seeded page-level split of the training projects (
partition=0.9,seed=42) - Page cap: 138
- `dh-unibe/image-text_nr-sr-vereinigte-bundesversammlung-xix`
- Training projects:
Nationalrat_01__Sitzung_–_06_11_1848,Nationalrat_02__Sitzung_–_07_11_1848,Nationalrat_03__Sitzung_–_08_11_1848,Nationalrat_04__Sitzung_–_09_11_1848,Nationalrat_05__Sitzung_–_10_11_1848,Nationalrat_06__Sitzung_–_13_11_1848,Nationalrat_07__Sitzung_–_14_11_1848,Nationalrat_08__Sitzung_–_15_11_1848,Nationalrat_09__Sitzung_–_18_11_1848,Nationalrat_10__Sitzung_–_21_11_1848,Nationalrat_11__Sitzung_–_22_11_1848,Nationalrat_12__Sitzung_–_23_11_1848,Nationalrat_13__Sitzung_–_24_11_1848,Nationalrat_14__Sitzung_–_25_11_1848,Nationalrat_15__Sitzung_–_27_11_1848,Nationalrat_16__Sitzung_–_28_11_1848,Nationalrat_17__Sitzung_–_29_11_1848,Ständerat_01__Sitzung_–_06_11_1848,Ständerat_02__Sitzung_–_08_11_1848,Ständerat_03__Sitzung_–_09_11_1848,Ständerat_04__Sitzung_–_11_11_1848,Ständerat_05__Sitzung_–_14_11_1848,Ständerat_06__Sitzung_–_15_11_1848,Ständerat_07__Sitzung_–_21_11_1848,Ständerat_08__Sitzung_–_23_11_1848,Ständerat_09__Sitzung_–_24_11_1848,Ständerat_10__Sitzung_–_25_11_1848,Ständerat_11__Sitzung_–_27_11_1848,Ständerat_12__Sitzung_–_28_11_1848,Ständerat_13__Sitzung_–_28_11_1848,VBV_01__Sitzung_–_16_11_1848,VBV_02__Sitzung_–_17_11_1848,VBV_03__Sitzung_–_20_11_1848,VBV_04__Sitzung_–_24_11_1848,VBV_05__Sitzung_–_29_11_1848 - Evaluation: a seeded page-level split of the training projects (
partition=0.9,seed=42) - Page cap: 52
- `dh-unibe/image-text_kurrent-xix`
- Training projects:
TRAIN_CITlab_Bassermann_0_4,TRAIN_CITlab_Bassermann_Manuscripts,TRAIN_CITlab_Bassermann_Manuscripts_0_2,TRAIN_CITlab_Binder_Kochbuch_2,TRAIN_CITlab_Escher_M1,TRAIN_CITlab_Gusbeth,TRAIN_CITlab_Handschriftliche_Archivquellen,TRAIN_CITlab_Konzilsprotkolle_B_Schwartz_5_2018,TRAIN_CITlab_MargareteSick,TRAIN_CITlab_MargareteSick20180731,TRAIN_CITlab_MargareteSick20180731a,TRAIN_CITlab_MÜLLER,TRAIN_CITlab_Protokoll_Hoftheater_1806,TRAIN_CITlab_Rehlen_1834_a,TRAIN_CITlab_RvE_Barlaam_HS_D_,TRAIN_CITlab_Steiner,TRAIN_CITlab_Suppes,TRAIN_CITlab_Suppes_3500,TRAIN_CITlab_Tagebuch_Arnold_v1,TRAIN_CITlab_umkc_Roland_M1,TRAIN_CITlab_umkc_Roland_M2,hufeland_privatbesitz_1829,nn_msgermqu2124_1827,nn_msgermqu2345_1827,parthey - Evaluation: held-out projects
TEST_CITlab_Bassermann_0_4,TEST_CITlab_Bassermann_Manuscripts,TEST_CITlab_Bassermann_Manuscripts_0_2,TEST_CITlab_Binder_Kochbuch_2,TEST_CITlab_Escher_M1,TEST_CITlab_Gusbeth,TEST_CITlab_Handschriftliche_Archivquellen,TEST_CITlab_Konzilsprotkolle_B_Schwartz_5_2018,TEST_CITlab_MargareteSick,TEST_CITlab_MargareteSick20180731,TEST_CITlab_MargareteSick20180731a,TEST_CITlab_MÜLLER,TEST_CITlab_Protokoll_Hoftheater_1806,TEST_CITlab_Rehlen_1834_a,TEST_CITlab_RvE_Barlaam_HS_D_,TEST_CITlab_Steiner,TEST_CITlab_Suppes,TEST_CITlab_Suppes_3500,TEST_CITlab_Tagebuch_Arnold_v1,TEST_CITlab_umkc_Roland_M1,TEST_CITlab_umkc_Roland_M2 - Page cap: 19808
Materialized from that selection: 26,186 pages, 964,472 transcribed lines, 966,748 training samples.
Trained with the instruction: Transcribe the handwritten text in this image exactly as written. — serving it with different wording is a silent distribution shift.
Hyperparameters
granularity: line
prompt: Transcribe the handwritten text in this image exactly as written.
load_in_4bit: false
lora_r: 64
lora_alpha: 128
lora_dropout: 0.05
target_modules:
- q_proj
- k_proj
- v_proj
- o_proj
- gate_proj
- up_proj
- down_proj
modules_to_save: []
epochs: 1
max_epochs: null
patience: 2
min_delta: 0.0001
batch_size: 16
accumulate_grad_batches: 1
lrate: 0.0002
lr_scheduler: cosine
warmup_ratio: 0.05
weight_decay: 0.0
max_grad_norm: 1.0
optim: paged_adamw_8bit
gradient_checkpointing: true
save_steps: 200
max_pixels: 262144
max_seq_len: 1024
min_train_chars: 0
eval_samples: 200
max_new_tokens: null
seed: 42
workers: 8
device: cuda:0
wandb_run: nullProvenance
metadata.json in this repo is the record the trainer wrote, verbatim: the full request, the parsed metrics and the job id.
Using it
This is a LoRA adapter, not a full model — it needs its base:
from peft import PeftModel
from transformers import AutoModelForImageTextToText, AutoProcessor
base = AutoModelForImageTextToText.from_pretrained('Qwen/Qwen3-VL-4B-Instruct')
model = PeftModel.from_pretrained(base, 'dh-unibe/qwen3vl-german-xix-v2')
processor = AutoProcessor.from_pretrained('dh-unibe/qwen3vl-german-xix-v2', trust_remote_code=True)vLLM 0.11 will not serve it as an adapter (it refuses LoRA on the vision tower), so serving means merging it into the base first — scripts/merge_loras.py in serving-atr-inference does that.
Notes
HELD-OUT RESULT. On the published benchmark 'Handwritten Text Recognition Test Set: Minutes of the Swiss Federal Council (1848-1903)' (Hodel & Schoch 2021, Zenodo, https://doi.org/10.5281/zenodo.4746342; private HF mirror dh-unibe/image-textfederal-minutes-testset), all 2,751 lines, no document shared with training: CER 0.0765, WER 0.2458, lengthratio 1.0019.
WHAT THIS REPLACES. qwen3vl-german-xix-v1 scored CER 0.2551 on the same lines and wrote little more than the first word on 507 of them (18.4%). v2 collapses on 0. The only difference is the corpus: v1 was built before the PageXML fix 33f55fc (#125), which had truncated 6.35x the characters of nr-sr-vereinigte-bundesversammlung-xix and 4.54x of parlamentsdienste-protokolle to their first word. Same four repositories, same seed, same page-level split.
THE CER IN THE TABLE ABOVE (0.0533) IS THE VALIDATION SPLIT, not the benchmark, and it is NOT comparable with v1's 0.0100: v1 was scored on the first 200 lines of val.jsonl (five in-domain pages), v2 on the stratified draw introduced in #120. Quote 0.0765 as this model's accuracy on unseen hands, and 0.0533 only as its score on its own corpus.
The remaining errors are 5,499 substitutions against 1,538 missing characters - misreadings rather than lost text.
