CoolFace
Modelpublic

ratlinghitman/qwen3vl-4b-receipt-extraction

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes11downloads
Model Card

Qwen3-VL-4B Receipt Extraction (bf16)

Reads a receipt image and returns it as structured JSON. Full-precision master: the merged bf16 model every other artifact in this project is derived from.

At a glance

Base model`Qwen/Qwen3-VL-4B-Instruct`
Fine-tuningLoRA r=16, language tower only
Dataset`naver-clova-ix/cord-v2`
Formatbf16 safetensors
Size on disk8.3 GB
Runtimetransformers

Intended use

Turning a photo or scan of a retail receipt into structured JSON, for bookkeeping and expense pipelines, or as a starting point for research on document understanding.

Out of scope

Invoices, handwritten notes, and non-receipt images. The model was never trained on them. It also should not be trusted to make financial decisions on its own; treat the output as an extraction to be reviewed.

How to use

The model expects one image and this exact prompt, which is the prompt it was trained with:

Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!
python
import torch
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor

REPO = "<your-hf-username>/qwen3vl-4b-receipt-extraction"
model = Qwen3VLForConditionalGeneration.from_pretrained(REPO, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(REPO, max_pixels=1600 * 28 * 28)

messages = [{"role": "user", "content": [
    {"type": "image", "image": "receipt.jpg"},
    {"type": "text", "text": "Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!"},
]}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt",
).to(model.device)

out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Output

A single JSON object with three top-level keys: menu (the line items), sub_total and total. Every value is a string, copied from the receipt.

json
{"menu": [{"nm": "Es Teh Manis", "cnt": "2", "price": "8,000"}],
 "sub_total": {"subtotal_price": "8,000"},
 "total": {"total_price": "8,000", "cashprice": "10,000", "changeprice": "2,000"}}

Evaluation

CORD-v2 test split, 100 receipts, greedy decoding, images capped at 1600 patches.

field_f1nTEDjson_validityexact_match
0.88620.91050.99000.4500

Measured in results/finetune-result.json (eval_script.py, plain decode).

  • —field_f1 flattens the predicted and gold JSON into (field-path, value) leaf pairs and scores precision/recall over them. It asks whether the right values landed at the right paths.
  • —nTED (normalized tree-edit distance, following the CORD paper) compares the two JSON trees directly. It catches structural mistakes, like wrong nesting or a missing key, that leaf-flattening misses.

Neither subsumes the other, which is why both are here.

Training

Fine-tuned from Qwen/Qwen3-VL-4B-Instruct with LoRA on `naver-clova-ix/cord-v2` (800 train / 100 validation / 100 test receipts). Labels were losslessly normalized before training.

LoRA rank / alpha / dropout16 / 32 / 0.05
Adapted moduleslanguage tower only (q,k,v,o,gate,up,down_proj); vision tower frozen
Precisionbf16
Learning rate1e-4
Batch size2 per device x 4 gradient accumulation
Epochs20
Image budget1600 patches of 28x28

Training ran well past convergence on purpose, so there was a wide range of adapters to pick from. Validation loss and the task metrics peak at different times here: loss starts rising around epoch 3, while the best fieldf1 and nTED land at epoch 9 or later. The adapter published here is the one that scored best on the **validation** split by fieldf1, and it was then scored once on the held-out test split. Early-stopping on loss would have picked a noticeably worse one.

Limitations

  • —The test split is 100 receipts. Gaps of a point or two in field_f1 are inside the noise band and should be read as ties.
  • —CORD-v2 is photographed Indonesian retail and restaurant receipts. Expect worse results on other layouts, languages, or document types (invoices, handwriting).
  • —Field names mirror the raw CORD keys (nm, cnt, unitprice), not friendly names.
  • —Values are transcribed text, not validated arithmetic. Nothing checks that the line items sum to the total. Do not use the output for financial decisions without review.
  • —Schema-constrained decoding was tried and scored far worse than plain decoding (field_f1 falls roughly 35 points), so plain greedy decoding is what these numbers use and what is recommended.
  • —Any timing or throughput figure depends heavily on hardware. Measure on the machine you intend to deploy on.

License

apache-2.0, inherited from the base model. The CORD-v2 dataset carries its own terms; see its dataset card.