Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF
Kimi-Linear-48B-A3B-Instruct — APEX GGUF
MoE-aware, mixed-precision APEX quantizations of moonshotai/Kimi-Linear-48B-A3B-Instruct — 48B total / ~3B active, a hybrid linear-attention MoE: most layers use KDA (Kimi Delta Attention, gated-delta linear attention), a few use full MLA attention, over a 256-routed + 1-shared expert FFN.
To my knowledge this is the first APEX quant of a linear-attention hybrid MoE. APEX assigns precision per tensor role and per layer instead of uniformly; here that meant teaching the recipe about tensor families the stock generator doesn't know (see Method).
⚠️ Set --ctx-size explicitly — do not run this model with defaults
Kimi-Linear was trained with up to a 1,048,576-token (1M) context window. If you launch llama-cli / llama-server without an explicit `--ctx-size`, llama.cpp defaults the KV cache to the model's own trained context length, not a small sane default — an attempt to allocate a KV cache sized for a million tokens, which can consume very large amounts of memory and stall or crash a machine with limited RAM/VRAM.
Always pass `--ctx-size` sized to what you actually need, e.g. --ctx-size 8192 for typical chat/tool-use. Only reach for six-figure-plus context sizes if you have the RAM/VRAM to back it.
Results
Perplexity on wikitext-2-raw (test, 200×512-token windows), llama-perplexity.
All three tiers land within ~0.3% of the bf16 reference. The imatrix-guided i-quality tier (IQ4_XS mid experts, 29 GB) is the smallest here and is available as Kimi-Linear-48B-A3B-Instruct-APEX-i-quality.gguf: at 4.94 BPW it trades ~0.34% perplexity for another ~4 GB off balanced.
From a 92 GB bf16 baseline → 33 GB (~2.8× smaller), and it runs on a 128 GB unified-memory box (fits with full GPU offload). Coherent on general and factual prompts.
The `balanced` and `handroll` tiers were built without an imatrix (Q6K/Q5K experts, Q80 shared, Q6K attention — none of which require importance data). The newer i-quality tier is imatrix-guided (IQ4_XS mid experts). Deeper lower-bit "I-tier" variants (IQ3/IQ2) would also need an imatrix and are not included here.
Note on the two tiers (a null result)
The hand-roll tier pins the KDA recurrence tensors (ssm_conv1d_*, ssm_f/g_*, ssm_beta) to Q8_0 instead of Q6_K, testing whether protecting the linear-attention state preserves quality. It doesn't — PPL is identical within noise (7.382 vs 7.377), at the same size (the ssm tensors are tiny next to the experts). Use `balanced`. The hand-roll is kept only to document the experiment.
Which file
- APEX-balanced — recommended. Q6K/Q5K experts on a layer-depth gradient, Q80 shared experts, Q6K attention + KDA tensors.
- APEX-i-quality — imatrix-guided IQ4_XS mid experts (29 GB); the smallest tier here (PPL 7.399, +0.34% vs bf16). Try it when you want a few GB over
balanced. - APEX-handroll — experimental (see null-result note below); not recommended.
Usage (llama.cpp)
llama-cli -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --ctx-size 8192 -p "Hello"
llama-server -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --ctx-size 8192 --host 0.0.0.0 --port 8080Requires a llama.cpp build supporting the kimi_linear architecture and the kimi-k2 pre-tokenizer.
Method
APEX is a bit-allocation recipe over stock llama-quantize --tensor-type-file. Kimi-Linear needed two tensor families the stock APEX generator doesn't emit:
- MLA (full-attention layers):
attn_kv_a_mqa,attn_k_b,attn_v_b - KDA (linear-attention layers):
ssm_conv1d_{k,q,v},ssm_f_a/f_b,ssm_g_a/g_b,ssm_beta(norms/1-D state kept F32)
plus a dense layer 0 (--dense-layers 1). The expert intermediate dim is 2048 (256-divisible), so no IQ4NL workaround was needed. Config generation + patching: see [`REPRODUCE.md`](REPRODUCE.md), `patchkimi_config.py, and configs/`.
Baseline: quantized from bartowski's bf16 GGUF.
Attribution & licenses
All MIT; see `LICENSE` and `NOTICE`.
- Base: Moonshot AI (@moonshotai) — Kimi-Linear-48B-A3B-Instruct (MIT)
- bf16 GGUF: bartowski (@bartowski) — source
- Engine: llama.cpp (@ggml-org · github) (MIT)
- APEX: Ettore Di Giacinto / LocalAI (@mudler) — localai-org/apex-quant (MIT)
Unofficial community quantization; not affiliated with or endorsed by Moonshot AI.
