CoolFace
Modelpublic

isaacmg/qwen3-vl-8b-hebrew-v18a-ckpt

sourceHugging Faceapache-2.0updated 22d agoView on Hugging Face
0likes437downloads
Model Card

Qwen3-VL-8B Hebrew — v1.8a checkpoints (vision-frozen control)

➡️ Unless you are studying this series' training history, use the flagship instead: [qwen3-vl-8b-hebrew-v19a-ckpt](https://huggingface.co/isaacmg/qwen3-vl-8b-hebrew-v19a-ckpt).

Status: superseded — A/B control arm. v1.8 was a controlled experiment on whether training the vision tower matters for handwriting: v1.8a kept the vision LoRA frozen (the accidental status quo of v1–v1.7, here made deliberate), while its twin v1.8b trained it. v1.8b won the comparison and became the ancestor of every later checkpoint. Keep this repo for reproducing the ablation; do not use it for inference.

The series at a glance

One line per generation — which checkpoint to use and which are historical:

versionrepostatuscanonical revisionheadline (benchmark)
v1.9av19a-ckptFLAGSHIP — use thisstep 1300 43e21bd7F1 0.816 / CER 0.216 (Genizah religious-140); F1 0.862 / CER 0.196 (frozen PGP-131)
v1.9bv19b-ckptexperimental control — not adoptedstep 1300 f8618c28merger-LoRA ablation study; loss ≡ v1.9a
v1.8bv18b-ckptsuperseded; warm-start ancestor of all v1.9step 700 c80313f8first arm with a live vision-tower LoRA
v1.8av18a-ckptsuperseded A/B control (vision frozen)step 700control arm for the v1.8 vision experiment
v1.7v17-ckptsupersededstep 800first Genizah-handwriting generation; best VLM on both corpora at its era's benchmarks (Aug 2026)
v1.6v16-ckptsuperseded; Talmud-print referencestep 1000 (official); step 1100 b9f47f32 (v1.7 warm start)Talmud page CER: gemara 0.090 / rashi 0.047 / tosafot 0.099 — Rashi-script 6.7× better than the best closed model we tested (0.315)
v1.5rashi-ckptsupersededfinal e050aca8 (step 2000)proved pure Rashi-glyph perception: CER 0.018 on unmemorizable synthetic text
v1hebrew-ckptsuperseded (language-only LoRA)step 3800 e2f85ddefirst checkpoints to read Vilna gemara from pixels

\* CER over substantive attempts only, alignment-based scorer. Benchmarks differ across generations (Talmud print vs Genizah manuscripts) — compare within a row's named benchmark, not across rows.