CoolFace
Modelpublic

HaithamalWaisy/kimi-k3-slice-Q3K-8L.gguf

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes616downloads
Model Card

Kimi K3 — 8-Layer Q3_K Serving Slice (5.26 GB)

An 8-layer prefix slice of the merged Kimi K3 model, quantized to a Q3K trunk with Q80 embeddings/output. Runs on a 2019 laptop (Intel Core i5-8265U, 7.4 GB single-channel RAM) at ~2.5–2.6 tokens/s single-stream.

Derived from: kimi-k3-merged-2exp.gguf (MergeMoE merge of Kimi K3).

Paper: Kimi K3 Under Compute Constraints: Implementing MergeMoE to Take 14 Shards to 5GB on Consumer Hardware

Author: Haitham al-Waisy · Contact: haithamalwaisy@gmail.com · X/Twitter · YouTube

Code: https://github.com/Haitham-alwaisy/Kimi-K3-MergeMoE

Verified behaviors

  • —Loads cleanly, generates real tokens, exits cleanly
  • —~2.5–2.6 tokens/s single-stream decode, ~5.3 t/s aggregate (6 batched streams)
  • —218 tensors, 4.90 GiB (5.26 GB decimal)

Usage

bash
hf download HaithamalWaisy/kimi-k3-slice-Q3K-8L.gguf kimi-k3-slice-Q3K-8L.gguf --local-dir .

Caveats

  • —Not benchmarked for quality end-to-end (see paper Section 4.1)
  • —Merge erases expert routing; slice keeps only 8 of 93 layers

License

Apache 2.0 (inherited from Kimi K3).