RepublicOfKorokke/GLM-OCR-oQ8-fp16
395
GLM-OCR-oQ8-fp16
This model was quantized using oQ mixed-precision quantization.
float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3/M4 and for numerical stability.
This model was quantized using oQ mixed-precision quantization.
float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3/M4 and for numerical stability.