CoolFace
Modelpublic

wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3-v1

sourceHugging Facemitupdated 22d agoView on Hugging Face
3likes448downloads
Model Card

GLM-5.3 Flash EXL3 K3

Three-bit EXL3/MCG for every routed expert in `zai-org/GLM-5.3-Flash-BF16`, with everything else retained from the official checkpoint.

What moved—and what did not

  • —K3 EXL3/MCG: all 288 routed experts' gate_proj, up_proj, and down_proj tensors in target layers 3…44 plus checkpoint MTP layer 45.
  • —37,152 quantized projections in one target-plus-MTP calibration run.
  • —Attention, dense MLPs, shared experts, routers, norms, embeddings, vision, and all other tensors remain native precision.
  • —The release audit compared 1,618 native tensors / 18.01 GiB byte-for-byte against source revision f12e0fe1f6b2ea274c11a569582edfd99d993c5e.
  • —Packed checkpoint payload: 127.30 GiB across 16 safetensors shards.

Quantization used GPTQModel 0565af7ce20a93df9bbc0e5563d7c6f60916f41a, EXL3 MCG K3, seed 787, sigma_reg=0.025, automatic output-scale selection, and a fixed 1,426-record calibration corpus (sha256:4e569625d97865777da92167b8fbf6fabb4ab7adf55baa5d676c81ce8dd95244). Natural GLM router traffic supplied expert Hessians; the committed recovery contract covers any expert below the 1,024-route floor. Quantization provenance and the validation report ship with the model; the full error ledger is retained separately as internal quantization evidence.

Before upload, the fully materialized checkpoint is loaded directly by the two-GPU B12x/vLLM recipe with adaptive MTP5 and NVFP4 MLA. Publication is gated on ordinary generation plus five strict tool-call scenarios. The exact prompts, expected behavior, actual calls, and raw report ship in quantization/vllm-prompt-expected-actual.md and quantization/vllm-tool-eval.json.

Serving

This is intended for the optimized two-GPU GLM-5.3 recipe at `tpurtell/glm-5.3-flash-ext3-4-bit-2x-rtx`. That runtime carries the GLM/EXL3/MTP and B12x ports; this card does not claim stock-vLLM support.

Huge thanks to Z.ai for GLM-5.3 Flash, Brandon for the earlier K4 quant and qualification work that helped inform this run, MiaAI-Lab for the nearby dual-DGX-Spark reference, and the GPTQModel, ExLlamaV3, vLLM, and B12x contributors.

Audit identity

  • —Source index: sha256:e6007bd58fb7e07f9fe69544257ee2713f252ef5855bbf685b48c991d524ef0f
  • —One-shot plan: sha256:fdce3c2c0078e27b0fa8bdab92cbd13f6c0ab34fb8d3e836fc56b44afc4ec82c
  • —Validation: sha256:625dcdfc8a031506d1406804cc72b7373c600567e076220b756f9504bc0fd284