ratlinghitman/qwen3vl-4b-receipt-extraction
Qwen3-VL-4B Receipt Extraction (bf16)
Reads a receipt image and returns it as structured JSON. Full-precision master: the merged bf16 model every other artifact in this project is derived from.
At a glance
Intended use
Turning a photo or scan of a retail receipt into structured JSON, for bookkeeping and expense pipelines, or as a starting point for research on document understanding.
Out of scope
Invoices, handwritten notes, and non-receipt images. The model was never trained on them. It also should not be trusted to make financial decisions on its own; treat the output as an extraction to be reviewed.
How to use
The model expects one image and this exact prompt, which is the prompt it was trained with:
Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!import torch
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
REPO = "<your-hf-username>/qwen3vl-4b-receipt-extraction"
model = Qwen3VLForConditionalGeneration.from_pretrained(REPO, dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(REPO, max_pixels=1600 * 28 * 28)
messages = [{"role": "user", "content": [
{"type": "image", "image": "receipt.jpg"},
{"type": "text", "text": "Extract the receipt from the image into a structured JSON. Your output should contain ONLY correct JSON!"},
]}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt",
).to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))Output
A single JSON object with three top-level keys: menu (the line items), sub_total and total. Every value is a string, copied from the receipt.
{"menu": [{"nm": "Es Teh Manis", "cnt": "2", "price": "8,000"}],
"sub_total": {"subtotal_price": "8,000"},
"total": {"total_price": "8,000", "cashprice": "10,000", "changeprice": "2,000"}}Evaluation
CORD-v2 test split, 100 receipts, greedy decoding, images capped at 1600 patches.
Measured in results/finetune-result.json (eval_script.py, plain decode).
- field_f1 flattens the predicted and gold JSON into (field-path, value) leaf pairs and scores precision/recall over them. It asks whether the right values landed at the right paths.
- nTED (normalized tree-edit distance, following the CORD paper) compares the two JSON trees directly. It catches structural mistakes, like wrong nesting or a missing key, that leaf-flattening misses.
Neither subsumes the other, which is why both are here.
Training
Fine-tuned from Qwen/Qwen3-VL-4B-Instruct with LoRA on `naver-clova-ix/cord-v2` (800 train / 100 validation / 100 test receipts). Labels were losslessly normalized before training.
Training ran well past convergence on purpose, so there was a wide range of adapters to pick from. Validation loss and the task metrics peak at different times here: loss starts rising around epoch 3, while the best fieldf1 and nTED land at epoch 9 or later. The adapter published here is the one that scored best on the **validation** split by fieldf1, and it was then scored once on the held-out test split. Early-stopping on loss would have picked a noticeably worse one.
Limitations
- The test split is 100 receipts. Gaps of a point or two in field_f1 are inside the noise band and should be read as ties.
- CORD-v2 is photographed Indonesian retail and restaurant receipts. Expect worse results on other layouts, languages, or document types (invoices, handwriting).
- Field names mirror the raw CORD keys (
nm,cnt,unitprice), not friendly names. - Values are transcribed text, not validated arithmetic. Nothing checks that the line items sum to the total. Do not use the output for financial decisions without review.
- Schema-constrained decoding was tried and scored far worse than plain decoding (field_f1 falls roughly 35 points), so plain greedy decoding is what these numbers use and what is recommended.
- Any timing or throughput figure depends heavily on hardware. Measure on the machine you intend to deploy on.
License
apache-2.0, inherited from the base model. The CORD-v2 dataset carries its own terms; see its dataset card.
