kailasa-ngpt/gemma-4-31b-tamil-devanagari-ocr-bf16
Gemma 4 31B Tamil + Devanagari OCR — BF16
This is the primary-quality merged checkpoint produced by the workspace's Tamil and Devanagari OCR fine-tune. It starts from google/gemma-4-31B-it and merges the selected rank-32 language-layer LoRA into Google's original BF16 weights. The vision tower stayed frozen. Sanskrit is covered where it is written in Devanagari; this model has not been validated for other Sanskrit scripts such as Grantha.
Training
- Distillation dataset:
kailasa-ngpt/gemini-3.7-flash-ocr-26-aug-2026 - Target rows: 8,374 Devanagari and 4,066 Tamil; 68 Telugu rows excluded
- Balanced training stream: 16,748 examples through deterministic Tamil oversampling
- Optimized examples: 4,000 (1,000 steps, physical batch 4)
- QLoRA: NF4 load, rank/alpha 32/32, 410 language-layer targets
- Learning rate: 1e-4 cosine; best checkpoint at step 950
- Best clean-book validation loss: 0.407109946
- Sequence/image budgets: 8,192 total tokens and 1,120 vision tokens
- Training peak: 31.13 GiB allocated on one RTX PRO 6000 96GB
Evaluation
All text metrics preserve raw model Unicode and line boundaries. The primary metric is extended-grapheme-cluster CER (gCER).
The fine-tune beat the base on 66/80 paired pages. A 20,000-replicate, script-stratified bootstrap estimated 99.505% probability that its raw mean gCER is lower; the very wide interval is caused by unbounded CER on base-model repetition loops.
On the untouched public benchmark's 350 natural-text cases, this BF16 artifact scored 3.238% gCER (Tamil 5.512%, Devanagari 0.965%) and 53.14% exact match. Its aggregate gCER difference from the base was not statistically resolved on that short-text benchmark.
vLLM
From the workspace root:
scripts/serve_vllm.sh artifacts/gemma4-31b-tamil-devanagari-ocr-bf16
.venv-serve/bin/python scripts/ocr_vllm_client.py page.png --output page.txtThe client uses the exact training prompt, temperature-zero decoding, preserves output verbatim, and cancels long exact grapheme repetition cycles. Exit code 2 means the token limit was reached; exit code 3 means the repetition guard fired. The artifact pins the 1,120-token image budget in both model and processor configuration, so ordinary vLLM loading matches the fine-tuning input budget.
Limitations and terms
The source dataset card declares license other, and its labels are Gemini teacher outputs rather than human ground truth. It also documents teacher-side normalization. Review the source-corpus rights, dataset terms, Gemini terms, and the base model's Apache 2.0 license before redistribution or commercial deployment. OCR errors, including occasional autoregressive loops, remain; do not treat the output as a faithful scholarly edition without review.
