CoolFace
Modelpublic

malaiwah/GLM-5.2-SIQ-Fruit-fp8

sourceHugging Facemitupdated 22d agoView on Hugging Face
0likes113downloads
Model Card

fruit-fp8, block-scaled FP8 e4m3

A FIXTURE for rehearsing quant-fidelity-suite's candidate route, not a serving quantization. Produced by engines/tools/fp8_quantize.py from malaiwah/GLM-5.2-SIQ-Fruit-bf16@ef68013aa6e16453cf52b5b77647f72fbe258c3c: every attention/indexer/MLP/expert projection weight is fp8 e4m3 with an fp32 weight_scale_inv per 128x128 block (scale = block amax / 448, ceil-padded grid, partial last blocks kept), the same checkpoint form zai-org/GLM-5.3 ships; embeddings, head, norms, router and MTP glue stay bf16 and are listed in quantization_config.modules_to_not_convert.

8588 tensors quantized; 102 modules native.