TheDrainFlorist/Qwen3.8-Flash-Next-VQ-4.4bpw
card: requirements up front (mlx-vlm >= 0.6.16 for images; model_file loaders only) + vision-fix changelog
fix: reach the vision tower, without giving up mlx_lm text serving
card: full ladder re-measured on the corrected scorer (2026-09-15). Every VQ rung beats the affine rung at or above its size; earlier published KL/ppl figures came from a streamed scorer since found to disagree with a direct forward and are superseded.
card: re-measured KL ladder across the VQ rungs (corrected scorer, 2026-09-15)
card: MTP toggle is exo --mtp (verified live, acceptance 0.875-1.0); note sequential serving under MTP
Card: fix strikethrough rendering (~ -> ≈)
2026-09-09 runtime refresh: faster prefill kernels (bit-identical), self-contained bundle, MTP sidecar, updated card
2026-09-09 runtime refresh: faster prefill kernels (bit-identical), self-contained bundle, MTP sidecar, updated card
card: exo serving via the mtp-stage1 fork; MTP drafting is opt-in with EXO_MTP=1 (measured 23.7-25.5 tok/s, acceptance 0.82-0.88 on the 2.1bpw rung)
add video_preprocessor_config.json: image/video processor configs from the upstream base (required by the mlx_vlm/exo vision path)
add preprocessor_config.json: image/video processor configs from the upstream base (required by the mlx_vlm/exo vision path)
2026-09 refresh: +8% kernels (1-ULP equiv, zero measured quality delta), cache ceiling, MTP sidecar (opt-in), three-point measured memory figures on card
2026-09 refresh: +8% kernels (1-ULP equiv, zero measured quality delta), cache ceiling, MTP sidecar (opt-in), three-point measured memory figures on card
2026-09 refresh: +8% kernels (1-ULP equiv, zero measured quality delta), cache ceiling, MTP sidecar (opt-in), three-point measured memory figures on card
2026-09 refresh: +8% kernels (1-ULP equiv, zero measured quality delta), cache ceiling, MTP sidecar (opt-in), three-point measured memory figures on card
Document that this model needs mlx-lm PR #1788 (qwen4_exp) to run
card: MTP head not included
config: remove MTP graft marker
pull MTP graft — no runtime implements MTP; re-ships only as a working, measured feature
card: bf16 MTP graft
MTP graft: bf16 verbatim (replaces 8-bit; smoke-gated)
MTP graft: bf16 verbatim (replaces 8-bit; smoke-gated)
card: 8-bit MTP graft shipped
add 8-bit MTP graft (disk-only; inert under current runtimes)
add 8-bit MTP graft (disk-only; inert under current runtimes)
card: note the MTP head is not included (no runtime implements it)
d2/K256 + hot-6 K1024 mix, 94.1 GiB: KL 50.3, beats q6 at 43 GiB less
initial commit
