khanh2023/Qwen3.6-14B-A3B-FableVibes-mlx-q6
1457
Qwen3.6-14B-A3B-FableVibes-mlx-q6
MLX 6-bit quantization of **tvall43/Qwen3.6-14B-A3B-FableVibes**, for local inference on Apple Silicon.
Credit / original model
This repo is only a quantized MLX conversion. All credit for the model itself goes to the original author, [tvall43](https://huggingface.co/tvall43). Please see and cite the original model card.
The base is a REAP-pruned Qwen3.6-35B-A3B reduced to ~14B total / ~3B active (90 experts, 8 active), recovered with a QLoRA distill of Claude Fable 5 reasoning traces. It uses the Qwen3.5 hybrid architecture (GatedDeltaNet linear attention + full attention + MoE) and emits <think>...</think> reasoning.
What this conversion did
- Fused the routed-MoE experts from per-expert tensors (
experts.{i}.{gate,up,down}_proj) into mlx-lm's stackedexperts.gate_up_proj/experts.down_projformat. - Quantized to 6-bit, group size 64 with
mlx-lm. - ~10 GB; runs on a 16 GB Apple Silicon Mac.
Usage
uv run python -m mlx_lm generate \
--model khanh2023/Qwen3.6-14B-A3B-FableVibes-mlx-q6 \
--prompt "Solve: ..."Notes
- MoE sparsity (~3B active/token) makes decode fast (~46 tok/s on an M4) despite 14B total params.
- 6-bit preserves more exactness than q4 on strict reasoning, at ~10 GB (needs a raised Metal wired limit on 16 GB). A smaller
-mlx-q4variant is also available.
