CoolFace
Modelpublic

axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes236downloads
Model Card

GLM-5.3-Flash W4A16 NVFP4 GGUF

GGUF conversion of `axiomofmind/GLM-5.3-Flash-W4A16-NVFP4`, based on `zai-org/GLM-5.3-Flash-BF16`.

The main-model routed experts use W4A16 NVFP4 with group size 16. Attention, shared experts, routers, embeddings, the output head, MTP weights, and other retained tensors remain in BF16 or F32.

Files

FileDescriptionSize
GLM-5.3-Flash-W4A16-NVFP4-BF16attn-MTP.ggufText model with MTP weights204.1 GB

Requirements

A llama.cpp build with GLM5Next and NVFP4 GGUF support is required.

License

This model is distributed under the MIT License. Refer to the official model card for architecture details, usage guidance, and limitations.