CoolFace
Modelpublic

larshiakzemil/olmocr-farsi-lora

sourceHugging Faceapache-2.0updated 7d agoView on Hugging Face
0likes15downloads
Model Card

olmOCR Farsi LoRA

LoRA adapter only (not full 7B weights) for Persian / Farsi document OCR, fine-tuned on top of allenai/olmOCR-2-7B-1025 (Qwen2.5-VL).

Adapter size ≈ 39 MB. Load the base model from AllenAI, then attach this adapter with PEFT.

Training summary

ItemValue
Base modelallenai/olmOCR-2-7B-1025 (~7B)
MethodQLoRA (4-bit base + LoRA adapters)
LoRA rank / alphar=16, alpha=32, dropout 0.05
Target modulesq_proj, k_proj, v_proj, o_proj
Trainable params~10.1M (0.12% of 8.3B)
Hardware1× RTX 3090 24GB
Train / eval159,999 / 2,000 page pairs (80/20)
Batch4 (effective 4), image longest side 768
Learning rate2e-4 cosine, warmup 5%
Planned epochs100 (stopped early ~epoch 2.0)
Best checkpointcheckpoint-80000
Best eval_loss0.0487
Train runtime (to best)~63h

Eval loss curve

Stepeval_loss
10,0000.0963
20,0000.0841
30,0000.0736
40,0000.0665
50,0000.0586
60,0000.0559
70,0000.0497
80,0000.0487

Eval loss kept falling through step 80k; this upload is the best (and latest) saved adapter from that run.

Files

  • —adapter_model.safetensors — LoRA weights only
  • —adapter_config.json — PEFT config (base_model_name_or_path → allenai/olmOCR-2-7B-1025)

Optimizer / scheduler / training state were not uploaded.

Usage

python
import torch
from PIL import Image
from peft import PeftModel
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration

base_id = "allenai/olmOCR-2-7B-1025"
lora_id = "larshiakzemil/olmocr-farsi-lora"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32

processor = AutoProcessor.from_pretrained(base_id, trust_remote_code=True)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    base_id, torch_dtype=dtype, device_map=device, trust_remote_code=True
)
model = PeftModel.from_pretrained(model, lora_id)
model.eval()

image = Image.open("page.jpg").convert("RGB")
prompt = (
    "Attached is one page of a document that you must process. "
    "Just return the plain text representation of this document as if you were reading it naturally. "
    "Convert equations to LateX and tables to HTML.\n"
    "Return your output as markdown, with a front matter section on top specifying values for the "
    "primary_language, is_rotation_valid, rotation_correction, is_table, and is_diagram parameters."
)
messages = [{
    "role": "user",
    "content": [
        {"type": "image", "image": image},
        {"type": "text", "text": prompt},
    ],
}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt")
inputs = {k: v.to(device) if hasattr(v, "to") else v for k, v in inputs.items()}

with torch.inference_mode():
    out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.decode(out[0], skip_special_tokens=True))

Prefer the official olmOCR v4 YAML prompt (build_no_anchoring_v4_yaml_prompt) when using the olmOCR toolkit — that matches training.

Notes

  • —This is an adapter, not a merged full model. You still need the ~7B base from AllenAI.
  • —Trained for Persian line / page OCR; early-stopped around epoch 2 of a planned 100-epoch schedule.
  • —Companion full-SFT model (different architecture): `larshiakzemil/glm-ocr-farsi`.