CoolFace
Modelpublic

spicyneuron/GLM-5.1-MLX-3.6bit

sourceHugging Facemitupdated 4mo agoView on Hugging Face
1likes57downloads
Model Card

GLM 5.1 optimized to run comfortably on a Mac Studio M3 512. This is the balanced version. Alternatives: speed-first, quality-first

  • —A mixed-precision quant that balances speed, memory, and accuracy.
  • —3-bit baseline with important layers at 4, 8 and BF16.
  • —Fits into ~350 GB memory, leaving plenty of room to run parallel models (ex: Minimax M2.7, Qwen 3.6 35B).

Usage

sh
# Start server at http://localhost:8080/chat/completions
uvx --from mlx-lm mlx_lm.server \
  --host 127.0.0.1 \
  --port 8080 \
  --model spicyneuron/GLM-5.1-MLX-3.6bit

Benchmarks

metricbaa-ai/GLM-5.1-RAM-270GB-MLX2.9 bit3.6 bit (this model)4.5 bit
bpw3.1102.9063.6454.538
base memory269.303251.702315.648392.992
peak memory (1024/512)291.257272.358341.020424.067
prompt tok/s (1024)194.958 ± 0.075194.216 ± 0.167190.508 ± 0.880193.563 ± 0.094
gen tok/s (512)21.381 ± 0.05019.527 ± 0.03517.873 ± 0.15617.259 ± 0.032
kl mean\*0.686 ± 0.0540.268 ± 0.0090.117 ± 0.0040.048 ± 0.002
kl p95\*1.478 ± 0.0540.537 ± 0.0090.236 ± 0.0040.097 ± 0.002
perplexity4.780 ± 0.0204.118 ± 0.0163.945 ± 0.0163.920 ± 0.016
piqa0.776 ± 0.0100.794 ± 0.0090.820 ± 0.0170.814 ± 0.017

\* GLM 5.1 KL divergence calculated against the largest quant I could run locally (~495 GB), so real KL is higher.

Tested on a Mac Studio M3 Ultra with:

mlx_lm.kld --baseline-model path/to/mlx-full-precision
mlx_lm.perplexity --sequence-length 2048 --seed 123
mlx_lm.benchmark --prompt-tokens 1024 --generation-tokens 512 --num-trials 5
mlx_lm.evaluate --tasks piqa --seed 123 --num-shots 0 --limit 500

mlx_lm.kld is approximate, based on top_k not full logits. Here's the code.

Methodology

Quantized with a mlx-lm fork, drawing inspiration from Unsloth/AesSedai/ubergarm style mixed-precision GGUFs. MLX quantization options differ from llama.cpp, but the principles are the same:

  • —Sensitive layers like MoE routing, attention, and output embeddings get higher precision
  • —More tolerant layers like MoE experts get lower precision