CoolFace
Modelpublic

AesSedai/GLM-5.3-Flash-GGUF

sourceHugging Faceupdated 25d agoView on Hugging Face
6likes2.9kdownloads
Model Card

Notes

This repo contains specialized MoE-quants for zai-org/GLM-5.3-Flash-BF16. The idea being that given the huge size of the FFN tensors compared to the rest of the tensors in the model, it should be possible to achieve a better quality while keeping the overall size of the entire model smaller compared to a similar naive quantization. To that end, the quantization type default is kept in high quality and the FFN UP + FFN GATE tensors are quanted down along with the FFN DOWN tensors.

QuantSizeMixturePPL1-(Mean PPL(Q)/PPL(base))KLD
Q5KM224.28 GiB (6.01 BPW)Q80 / Q5K / Q5K / Q6K3.589877 ± 0.019865+0.5529%0.027859 ± 0.000207
Q4KM188.10 GiB (5.04 BPW)Q80 / Q4K / Q4K / Q5K3.635356 ± 0.020204+1.8267%0.050181 ± 0.000333
IQ4_XS148.24 GiB (3.97 BPW)Q80 / IQ3S / IQ3S / IQ4XS3.819227 ± 0.021423+6.9770%0.117358 ± 0.000727
IQ3_S116.14 GiB (3.11 BPW)Q6K / IQ2S / IQ2S / IQ3S4.387061 ± 0.025595+22.8821%0.283438 ± 0.001596
IQ2_S105.81 GiB (2.83 BPW)Q6K / IQ2XS / IQ2XS / IQ3XXS4.761384 ± 0.028305+33.3669%0.375406 ± 0.001984

[image] [image]