CoolFace
Modelpublic

Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes2.7kdownloads
Model Card

Kimi-Linear-48B-A3B-Instruct — APEX GGUF

MoE-aware, mixed-precision APEX quantizations of moonshotai/Kimi-Linear-48B-A3B-Instruct — 48B total / ~3B active, a hybrid linear-attention MoE: most layers use KDA (Kimi Delta Attention, gated-delta linear attention), a few use full MLA attention, over a 256-routed + 1-shared expert FFN.

To my knowledge this is the first APEX quant of a linear-attention hybrid MoE. APEX assigns precision per tensor role and per layer instead of uniformly; here that meant teaching the recipe about tensor families the stock generator doesn't know (see Method).

⚠️ Set --ctx-size explicitly — do not run this model with defaults

Kimi-Linear was trained with up to a 1,048,576-token (1M) context window. If you launch llama-cli / llama-server without an explicit `--ctx-size`, llama.cpp defaults the KV cache to the model's own trained context length, not a small sane default — an attempt to allocate a KV cache sized for a million tokens, which can consume very large amounts of memory and stall or crash a machine with limited RAM/VRAM.

Always pass `--ctx-size` sized to what you actually need, e.g. --ctx-size 8192 for typical chat/tool-use. Only reach for six-figure-plus context sizes if you have the RAM/VRAM to back it.

Results

Perplexity on wikitext-2-raw (test, 200×512-token windows), llama-perplexity.

FileSizeBPWPPLΔ vs bf16
bf16 (reference)92 GB16.07.374
APEX-balanced (no-imatrix)33 GB5.727.377+0.04%
APEX-handroll (ssm@Q8_0, no-imatrix)33 GB5.727.382+0.11%
APEX-i-quality (imatrix, IQ4_XS mid experts)29 GB4.947.399+0.34%

All three tiers land within ~0.3% of the bf16 reference. The imatrix-guided i-quality tier (IQ4_XS mid experts, 29 GB) is the smallest here and is available as Kimi-Linear-48B-A3B-Instruct-APEX-i-quality.gguf: at 4.94 BPW it trades ~0.34% perplexity for another ~4 GB off balanced.

From a 92 GB bf16 baseline → 33 GB (~2.8× smaller), and it runs on a 128 GB unified-memory box (fits with full GPU offload). Coherent on general and factual prompts.

The `balanced` and `handroll` tiers were built without an imatrix (Q6K/Q5K experts, Q80 shared, Q6K attention — none of which require importance data). The newer i-quality tier is imatrix-guided (IQ4_XS mid experts). Deeper lower-bit "I-tier" variants (IQ3/IQ2) would also need an imatrix and are not included here.

Note on the two tiers (a null result)

The hand-roll tier pins the KDA recurrence tensors (ssm_conv1d_*, ssm_f/g_*, ssm_beta) to Q8_0 instead of Q6_K, testing whether protecting the linear-attention state preserves quality. It doesn't — PPL is identical within noise (7.382 vs 7.377), at the same size (the ssm tensors are tiny next to the experts). Use `balanced`. The hand-roll is kept only to document the experiment.

Which file

  • APEX-balanced — recommended. Q6K/Q5K experts on a layer-depth gradient, Q80 shared experts, Q6K attention + KDA tensors.
  • APEX-i-quality — imatrix-guided IQ4_XS mid experts (29 GB); the smallest tier here (PPL 7.399, +0.34% vs bf16). Try it when you want a few GB over balanced.
  • APEX-handroll — experimental (see null-result note below); not recommended.

Usage (llama.cpp)

bash
llama-cli   -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --ctx-size 8192 -p "Hello"
llama-server -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --ctx-size 8192 --host 0.0.0.0 --port 8080

Requires a llama.cpp build supporting the kimi_linear architecture and the kimi-k2 pre-tokenizer.

Method

APEX is a bit-allocation recipe over stock llama-quantize --tensor-type-file. Kimi-Linear needed two tensor families the stock APEX generator doesn't emit:

  • MLA (full-attention layers): attn_kv_a_mqa, attn_k_b, attn_v_b
  • KDA (linear-attention layers): ssm_conv1d_{k,q,v}, ssm_f_a/f_b, ssm_g_a/g_b, ssm_beta (norms/1-D state kept F32)

plus a dense layer 0 (--dense-layers 1). The expert intermediate dim is 2048 (256-divisible), so no IQ4NL workaround was needed. Config generation + patching: see [`REPRODUCE.md`](REPRODUCE.md), `patchkimi_config.py, and configs/`.

Baseline: quantized from bartowski's bf16 GGUF.

Attribution & licenses

All MIT; see `LICENSE` and `NOTICE`.

Unofficial community quantization; not affiliated with or endorsed by Moonshot AI.