CoolFace
Modelpublic

RepublicOfKorokke/GLM-OCR-oQ8-fp16

sourceHugging Facemitupdated 5mo agoView on Hugging Face
3likes95downloads
Model Card

GLM-OCR-oQ8-fp16

This model was quantized using oQ mixed-precision quantization.

float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3/M4 and for numerical stability.

Benchmark (on M1 Max)

Model VariantPP (Tokens per second)TG (Tokens per second)
Original (bf16)4,684104.8
oQ8-fp163,80699.0