axiomofmind/GLM-5.3-Flash-W4A16-NVFP4-GGUF
0236
GLM-5.3-Flash W4A16 NVFP4 GGUF
GGUF conversion of `axiomofmind/GLM-5.3-Flash-W4A16-NVFP4`, based on `zai-org/GLM-5.3-Flash-BF16`.
The main-model routed experts use W4A16 NVFP4 with group size 16. Attention, shared experts, routers, embeddings, the output head, MTP weights, and other retained tensors remain in BF16 or F32.
Files
Requirements
A llama.cpp build with GLM5Next and NVFP4 GGUF support is required.
License
This model is distributed under the MIT License. Refer to the official model card for architecture details, usage guidance, and limitations.
