CoolFace
Modelpublic

Abiray/OvisOCR2-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
28likes5.4kdownloads
Model Card

OvisOCR2 - GGUF Quantizations

<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/658a8a837959448ef5500ce5/vRCIu5QD8VuIJolkC_ZHQ.png" alt="Ovis" width="30%" /> </p>

This repository contains GGUF format quantizations of OvisOCR2, a compact 0.8B end-to-end model for page-level document parsing. The original model was developed by ATH-MaaS by post-training Qwen3.5-0.8B to parse full document pages directly into clean Markdown (including LaTeX formulas, HTML tables, and layout components).

OvisOCR2 establishes a new state-of-the-art for compact document understanding, scoring 96.58 on OmniDocBench v1.6 and outperforming traditional, multi-stage layout analysis pipelines.


Available Files

Main Text Models

File NamePrecision / QuantizationFile SizeDescription
OvisOCR2-F16.gguf16-bit Float1.52 GBBaseline unquantized model
OvisOCR2-BF16.gguf16-bit Brain Float1.52 GBNative weight precision
OvisOCR2-Q8_0.gguf8-bit812 MBNear-identical precision to F16
OvisOCR2-Q6_K.gguf6-bit630 MBExcellent balance of size and accuracy
OvisOCR2-Q5_K_M.gguf5-bit (Medium)578 MBRecommended for low-resource deployment
OvisOCR2-Q5_K_S.gguf5-bit (Small)564 MBHighly optimized 5-bit layout
OvisOCR2-Q4_K_M.gguf4-bit (Medium)529 MBStandard 4-bit quantization
OvisOCR2-Q4_K_S.gguf4-bit (Small)505 MBLightweight 4-bit footprint
OvisOCR2-Q3_K_M.gguf3-bit (Medium)466 MBMaximum compression ratio

Multimodal Projectors (mmproj)

Note: Because OvisOCR2 is a vision-language model, you must download one of these image processing units alongside your choice of the text models listed above.

  • mmproj-F32.gguf (402 MB) - Unquantized full precision projector.
  • mmproj-F16.gguf (205 MB) - Recommended standard performance/size option.
  • mmproj-BF16.gguf (207 MB) - Target alternative precision layout.

Inference Guide (llama.cpp)

To run multimodal OCR tasks using these GGUF files, you need to use the llama-minicpmv-cli or llama-llava-cli tool (depending on your build version of llama.cpp) to handle simultaneous image and text tokens.

Basic Command Line Example

bash
# Run parsing via llama.cpp cli tools
./llama-minicpmv-cli \
  -m OvisOCR2-Q5_K_M.gguf \
  --mmproj mmproj-F16.gguf \
  --image /path/to/your/document_page.jpg \
  -p "<|im_start|>user\nExtract all readable content from the image in natural human reading order and output the result as a single Markdown document. Format formulas as LaTeX. Format tables as HTML: <table>...</table>. Preserve the original text without translation.<|im_end|>\n<|im_start|>assistant\n" \
  -n 4096 \
  --temp 0.0