CoolFace
Modelpublic

destynova002/Qwen3.6-27B-mixed-3-4bit-3.85bpw-mlx

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
1likes169downloads
Model Card

Qwen3.6-27B — greedy sensitivity mixed 3/4-bit (3.85 bpw, MLX affine)

Mixed-precision quantization of Qwen3.6-27B targeting 24 GB Apple Silicon Macs: near-4-bit quality at 12 GB of weights.

Method

Sensitivity-based greedy bit allocation (weight-space, calibration-free):

  1. 1.For every 2D linear layer, compute the relative 3-bit quantization error rel_err3(W) = ‖W − dequant3(W)‖F / ‖W‖F.
  2. 2.Rank layers by error; assign 4-bit to embeddings/lm_head and to the most sensitive layers up to a 30 % parameter budget; everything else gets 3-bit.
  3. 3.Standard MLX affine quantization, group size 64 (loads in mlx-lm and any engine supporting per-layer mixed MLX affine quants).

Note: on this architecture (hybrid GatedDeltaNet/SSM), weight-space ranking outperforms activation-weighted (imatrix-style) ranking — the recurrent gate projections are systematically under-protected by input-activation proxies (measured: KL 0.065 vs 0.18–0.20 for three imatrix variants at equal bpw). Gradient-based methods (DWQ) are not applicable (no MLX backward for the SSM ops).

Quality (measured vs bf16)

metricvalue
bits per weight3.85
weights on disk12 GB
KL(bf16 ‖ quant), teacher-forced, 5 prompts (FR prose/code/math)0.065
reference pointsuniform 3-bit: 0.30 · mixed34 heuristic: 0.24 · 4-bit ≈ 0.02

Speed

Measured on a MacBook Pro M5 Max, 128 GB with saragossa (effective weight bandwidth ~470 GB/s on this workload):

modedecode
autoregressive, GPU-resident~35 tok/s
+ speculative MTP (greedy T=0, bundled head)~46 tok/s

Decode is memory-bandwidth-bound, so expect roughly proportional numbers on other chips: an M-series Pro has about half the Max's bandwidth (~17-23 tok/s here), a base M-series about a quarter. mlx-lm runs it too, without the MTP speedup.

Runtime footprint ≈ 16 GB with saragossa's single-copy unified-memory loader — comfortable on a 24 GB Mac with context headroom.

mtp.safetensors is the Qwen3.6 MTP (NextN) head, 4-bit, taken from mlx-community/Qwen3.6-27B-OptiQ-4bit and bundled so speculative decoding works directly in engines that support it (measured acceptance 0.81 on this backbone). Engines that don't will ignore it.

Usage

bash
# mlx-lm
mlx_lm.generate --model <this-repo> --prompt "Bonjour !"

# saragossa — brew install azerozero/tap/saragossa
# (MTP speculative decode engages automatically at T=0)
saragossa run <this-repo>

Tooling

The quantization tooling is public: `tools/quant/` — sensitivity-based bit allocation (calibration-free) plus the KL evaluation script used for the numbers above. The README there documents the exact commands that reproduce this repo.

License

Apache-2.0, same as the base model.