destynova002/Qwen3.6-27B-mixed-3-4bit-3.85bpw-mlx
Qwen3.6-27B — greedy sensitivity mixed 3/4-bit (3.85 bpw, MLX affine)
Mixed-precision quantization of Qwen3.6-27B targeting 24 GB Apple Silicon Macs: near-4-bit quality at 12 GB of weights.
Method
Sensitivity-based greedy bit allocation (weight-space, calibration-free):
- For every 2D linear layer, compute the relative 3-bit quantization error
rel_err3(W) = ‖W − dequant3(W)‖F / ‖W‖F. - Rank layers by error; assign 4-bit to embeddings/lm_head and to the most sensitive layers up to a 30 % parameter budget; everything else gets 3-bit.
- Standard MLX affine quantization, group size 64 (loads in
mlx-lmand any engine supporting per-layer mixed MLX affine quants).
Note: on this architecture (hybrid GatedDeltaNet/SSM), weight-space ranking outperforms activation-weighted (imatrix-style) ranking — the recurrent gate projections are systematically under-protected by input-activation proxies (measured: KL 0.065 vs 0.18–0.20 for three imatrix variants at equal bpw). Gradient-based methods (DWQ) are not applicable (no MLX backward for the SSM ops).
Quality (measured vs bf16)
Speed
Measured on a MacBook Pro M5 Max, 128 GB with saragossa (effective weight bandwidth ~470 GB/s on this workload):
Decode is memory-bandwidth-bound, so expect roughly proportional numbers on other chips: an M-series Pro has about half the Max's bandwidth (~17-23 tok/s here), a base M-series about a quarter. mlx-lm runs it too, without the MTP speedup.
Runtime footprint ≈ 16 GB with saragossa's single-copy unified-memory loader — comfortable on a 24 GB Mac with context headroom.
mtp.safetensors is the Qwen3.6 MTP (NextN) head, 4-bit, taken from mlx-community/Qwen3.6-27B-OptiQ-4bit and bundled so speculative decoding works directly in engines that support it (measured acceptance 0.81 on this backbone). Engines that don't will ignore it.
Usage
# mlx-lm
mlx_lm.generate --model <this-repo> --prompt "Bonjour !"
# saragossa — brew install azerozero/tap/saragossa
# (MTP speculative decode engages automatically at T=0)
saragossa run <this-repo>Tooling
The quantization tooling is public: `tools/quant/` — sensitivity-based bit allocation (calibration-free) plus the KL evaluation script used for the numbers above. The README there documents the exact commands that reproduce this repo.
License
Apache-2.0, same as the base model.
