CoolFace
Modelpublic

GreenBitAI/GLM-5.3-Flash-4bit-paged

sourceHugging Facemitupdated 6d agoView on Hugging Face
0likes838downloads
Model Card

GLM-5.3-Flash-4bit-paged

Expert-paged build of pipenetwork/GLM-5.3-Flash-MLX-4bit. The weights that are read a fraction at a time live in their own containers, so a machine loads what it needs rather than all of it.

filesizeholds
model.safetensors5.89 GiBresident weights
experts.bin159.47 GiBrouted experts
mtp/3.90 GiBdraft head, off by default

Total 169.27 GiB. Of that, 165.37 GiB is the source build, whose bytes moved into containers rather than being copied, and 3.90 GiB is the draft head, which no published build of this model carries.

python
from gbx_lm.utils import load
model, tokenizer = load("GreenBitAI/GLM-5.3-Flash-4bit-paged")

Where the weights fit they are filled from experts.bin and the model runs the stock path at stock speed; where they do not, they stream from disk. Reading the machine decides that, not a flag.

To override that: GBX_PAGING=off holds the experts resident.

Checked at build time, while the source checkpoint was still there to compare against:

  • PASS layer-wise vs resident — 42 layers x 2 draws exact, 60.93 GiB peak for this gate

Quantization, tokenizer, chat template and licence are unchanged from pipenetwork/GLM-5.3-Flash-MLX-4bit.

The draft head

The mtp/ folder carries the model's own multi-token prediction head, converted from zai-org/GLM-5.3-Flash, so speculative decoding works from this repository alone. It is off unless asked for:

bash
GBX_GLM53_MTP=on python -m gbx_lm.generate --model GreenBitAI/GLM-5.3-Flash-4bit-paged --max-tokens 256 --prompt "..."
GBX_GLM53_MTP=on python -m gbx_lm.fastapi_server --model GreenBitAI/GLM-5.3-Flash-4bit-paged

It grows with the context. Measured 2026-09-18 on a 512 GB Mac Studio (M3 Ultra), greedy, decode timed from the first token, both columns from the same build:

contexthead offhead onspeedupaccepted
11,31727.1 tok/s38.31.41x78%
2,85427.535.71.30x76%
71531.631.81.01x65%

A short context has little to go on, so the head proposes less well and the win is small; by 11k it accepts 78% of what it proposes. Through the server on chat-shaped questions it accepts 63-75%.

Every token the head proposes is verified by the model itself, so the reply is the model's own either way; the head only saves passes over the weights.