neopolita/Qwen3.6-23B-A3B-Niwaki-v2.3-4bit-mlx
Qwen3.6-23B-A3B-Niwaki-v2.3-4bit-mlx
Qwen3.6-35B-A3B pruned to 22.6B total / ~3.8B active parameters, stored at 4-bit — 3.0× less expert memory than the 8-bit reference — with the model's program-state tracking and few-shot reasoning intact, and lower perplexity than v2.2 at identical bytes. Loads with the shim shipped in this repo; everything else is stock `mlx_lm`.
Niwaki (庭木) are Japan's garden trees, sculpted by meticulous pruning so that every branch serves the form of the whole. This model applies that spirit to a Mixture-of-Experts.
A paper with the full method and measurements is coming soon.
Project page, with the releases and their measurements side by side: [niwakiai.com](https://niwakiai.com/)
What changed from v2.2
Same layout, same byte budget, same kept experts: layers 10–29 keep 64 of their 256 experts and route only over them, the other twenty layers are untouched. v2.3 changes only the recovery training that follows the pruning. That alone lowers perplexity by 2.4% on WikiText-2 and 4.8% on C4, holds the task average, lifts multi-hop recall from 0.50 to 0.69 — above the reference's 0.62 — and takes two-step arithmetic in generation from 0.95 to 0.98, with the tracking and distractor families unchanged at the reference's level on both evaluation seeds.
Benchmarks
Full evaluation protocol: WikiText-2 (145 × 2048-token windows) and C4 (256 × 2048-token windows) perplexity; complete lm-evaluation-harness suites for arc_easy, hellaswag, piqa, winogrande and boolq; gsm8k on 100 problems under lm-eval's default 5-shot prompt (strict match).
The Niwaki line so far
One row per generation at its flagship byte point, all under the protocol above (the first-generation models were re-measured for this table; their own cards carry older 500-sample screens).
The line reads left to right as a trade. v1's flagship spent the fewest bytes and, on the benchmarks above, still holds up (96.7% task retention, gsm8k 0.90); what it lost is the model's program-state tracking — the value a variable holds through two-hop aliasing or past a distractor line (0.82 against the reference's 0.95–0.99) and two-step arithmetic in generation (0.40 against 0.63) — the damage the later generations were built to avoid. v2 cut deeper and lost gsm8k as well; v2.1 bought both back for 1.5× v1's bytes; v2.2 kept v2.1's bytes and recovered 4–6% of perplexity from a better choice of kept experts; v2.3 keeps v2.2's layout and experts and takes another 2–5% from better recovery training, with multi-hop recall now above the reference's.
Against quantizing the intact model to the same bytes
The alternative to pruning is to keep every expert and quantize harder. The fairest such control we can build is the 8-bit reference with only its routed experts recast to 2-bit (group 64, the same quantizer as this repo) and the backbone untouched: 10.1 GB of expert storage (0.294×), 11% fewer bytes than this model, measured under the identical protocol — once as cast, and once after the same recovery training this model received. The last three columns are our program-state battery (generation accuracy, range over two evaluation seeds): the value a variable holds through two-hop aliasing, the same past a distractor line, and two-step arithmetic.
Read it as a split, not a win. At this byte point the intact model with 2-bit experts is the stronger artifact on perplexity and the task average, and after recovery training it matches the reference's task average at 0.294×. What it does not recover, trained or not, is program-state function: two-step arithmetic stays at 0.37–0.60 and two-hop tracking at 0.78–0.82, against 0.97–0.98 and 0.97 here — 2-bit noise on every expert breaks exact state carrying, and training does not restore it, whereas this model keeps those families at or above the reference's level. Pick by workload. The 2-bit sibling (0.213×) sits below the smallest byte point an intact 2-bit cast reaches in MLX (0.294× at group 64).
Evaluation methodology
Task average: the complete test sets of the five benchmarks scored by per-choice log-likelihood — arceasy, hellaswag, and piqa report normalized accuracy (`accnorm); winogrande and boolq report accuracy (acc`); the average is the unweighted mean of the five. Perplexity rows use the full window protocol above, identical windows for every row, loss in fp32, no chat template. gsm8k is lm-evaluation-harness's default 5-shot prompt with greedy decoding and strict answer matching, 100 problems. Generation diversity: an 8-prompt battery (code, reasoning, chat, creative; 600-token sampled generations) scored by bigram diversity (distinct bigrams / total bigrams; d2 is a diversity statistic, not a quality score — the reading is closeness to the reference's 0.63 / 0.51), every row measured with the same script, prompts and seeds, the reference alongside.
Model dimensions
Usage (MLX, Apple Silicon)
This checkpoint stores compact 64-expert banks for layers 10–29, which stock mlx-lm cannot express — load through the shim in this repo (it rebuilds those banks from config.json, then defers to mlx_lm.load):
# curl -O https://huggingface.co/neopolita/Qwen3.6-23B-A3B-Niwaki-v2.3-4bit-mlx/resolve/main/niwaki_v2_load.py
from niwaki_v2_load import load
from mlx_lm import generate
model, tokenizer = load("neopolita/Qwen3.6-23B-A3B-Niwaki-v2.3-4bit-mlx")
print(generate(model, tokenizer, prompt="...", max_tokens=256))GGUF builds for llama.cpp (through the neopolita-llama.cpp fork, which reads this model's per-layer expert counts): Qwen3.6-23B-A3B-Niwaki-v2.3-4bit-gguf (UD-Q3K recommended + Q4KM).
Notes and limitations
- Compression trades quality: this model retains ~97% of the reference's task average and ~86% of its gsm8k score. Choose the family member that fits your memory budget.
- Evaluated text-only on English-web-heavy data; the base model's biases are inherited and rare-domain behaviour is less tested.
- No latency claim: all 40 layers execute with the reference's active parameter count. This is a memory artifact.
- The chat template works as on the base model (thinking on or off). In raw completion mode the model continues few-shot patterns like a base model does; evaluate with stop sequences.
- Part of the Niwaki family: Qwen3.6-23B-A3B-Niwaki-v2.3-2bit-mlx (sibling), the v2.2 models 4-bit / 2-bit, the v2.1 models 4-bit / 2-bit, the v2 models 4-bit / 2-bit, and the first-generation 27B / 19B / 11B models.
Base model by the Qwen team (Apache 2.0); 8-bit MLX conversion by mlx-community; compression and distillation by the Niwaki project, 2026-09.
