kernelpool/GLM-5.3-5bit-UVMAX
1998
kernelpool/GLM-5.3-5bit-UVMAX
Mixed-precision (UVMAX) quantization of zai-org/GLM-5.3-BF16, converted with mlx-lm from the bf16 release.
What is UVMAX?
UVMAX assigns bit widths per tensor class instead of quantizing uniformly.
Quality
Teacher-forced against the bf16 release on identical tokens, 48 windows of 1025 tokens. KLD is KL(bf16 ‖ quant) over the full output distribution and is corpus-specific.
UVMAX (4.67 bits/weight)
Uniform 4-bit (4.50 bits/weight)
Use with mlx
Requires a recent mlx-lm with GLM-5.3 (glm_moe_dsa) support.
The model fits on a single 512 GB machine:
mlx_lm.server --model kernelpool/GLM-5.3-5bit-UVMAXSampling follows the base model: temperature 1.0, top-p 0.95.
