CoolFace
Modelpublic

objectai/obj_v1

sourceHugging Faceapache-2.0updated 22d agoView on Hugging Face
2likes541downloads
Model Card

obj_v1

Vision-language model fine-tuned for structured data extraction from Indian financial documents. Give it a page image and a JSON schema; it returns the schema filled in from what is on the page. A 4B vision-language model, LoRA fine-tuned and merged

Authors

<p align="left"> <!-- <a href="https://www.linkedin.com/in/ahmedzaweel/"> <img src="https://img.shields.io/badge/LinkedIn-Ahmed%20Zaweel-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Ahmed Zaweel on LinkedIn" /> </a> --> &nbsp;&nbsp; <a href="https://www.linkedin.com/in/rachit-kumar-b41299228/"> <img src="https://img.shields.io/badge/LinkedIn-Rachit%20Kumar-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Rachit Kumar on LinkedIn" /> </a> &nbsp;&nbsp; <a href="https://www.linkedin.com/in/ahmedzaweel/"> <img src="https://img.shields.io/badge/LinkedIn-Ahmed%20Zaweel-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Ahmed Zaweel on LinkedIn" /> </a> </p>

Serving with vLLM

bash
vllm serve objectai/obj_v1 \
  --served-model-name obj_v1 \
  --max-model-len 16384 \
  --limit-mm-per-prompt '{"image":1}' \
  --mm-processor-kwargs '{"max_pixels":1003520}' \
  --trust-remote-code

max_pixels is 1280x28x28, the resolution the model was trained at. Raising it wastes KV cache; lowering it makes small print unreadable.

Calling it

The server is OpenAI-compatible, so an ordinary chat completion works:

python
import base64, json, openai
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
image = base64.b64encode(open("cheque.jpg", "rb").read()).decode()
schema = {"cheque_details": {"amount": "number", "payee": "string",
                             "date": "string", "cheque_number": "string"}}
response = client.chat.completions.create(
    model="obj_v1",
    temperature=0.0,
    max_tokens=8192,
    messages=[
        {"role": "system", "content":
            "You are a document data extraction model. "
            "Extract only values present in the document. "
            "Use null for fields that are absent or illegible. "
            "Output a single compact JSON object matching the requested schema. "
            "No prose, no markdown, no explanation."},
        {"role": "user", "content": [
            {"type": "image_url",
             "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
            {"type": "text",
             "text": f"document_type: cheque\nschema: {json.dumps(schema)}"},
        ]},
    ],
)
print(response.choices[0].message.content)

Prompt format

Match training or accuracy drops. The system prompt above is verbatim, and the user turn is the image followed by exactly two lines:

document_type: <type>
schema: <compact json>

Set temperature=0.0 so the same page yields the same answer.

Requirements

Weights8.9 GB (bf16)
VRAM16 GB minimum, 24 GB comfortable
Precisionbf16 (Ampere or newer; use fp16 below that)
Context16384 covers the longest documents

Runs on an L4, A10G, L40S, A100 or RTX 4090. On a T4 add --dtype float16.

Output

Compact JSON matching the requested schema. Fields absent from the page come back null rather than guessed. Values found on the page that the schema did not ask for are placed under extras when that key is included in the schema.

Limitations

  • —Trained on Indian financial documents; other domains and layouts are untested.
  • —Handwriting is the weakest case, particularly digits at low resolution.
  • —The model does not verify its own arithmetic. Totals that must reconcile should be checked by the caller.

License

Apache 2.0. Fine-tuned from Qwen3-VL-4B-Instruct, which is Apache 2.0.