CoolFace
Modelpublic

KeinNiemand/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated-IK_GGUF

sourceHugging Faceupdated 4mo agoView on Hugging Face
4likes368downloads
Model Card

Qwopus3.5 122B A10B Kimi-K2.6 Distill Healed Abliterated - Custom GGUF Quantizations

CRITICAL COMPATIBILITY WARNING

These are `iqk` format quantizations and are EXCLUSIVE to the `ik_llama.cpp` fork.

They will NOT work on mainline llama.cpp, standard LM Studio, standard Text Generation WebUI, or KoboldCPP.

You must compile and run this using ikawrakow's llama.cpp fork, or a UI where you have manually swapped the backend to an ik_llama.cpp build.


This repository contains custom, mixed-precision ik_llama.cpp GGUF quantizations for OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated, a Kimi-K2.6 distilled, healed, abliterated Qwen3.5 122B A10B MoE model.

These quants use different precision levels for different layer types, keeping attention, SSM, shared expert, output, and MTP/NextN tensors at higher precision while compressing the routed experts, which make up the bulk of the model's size.

⚠️ Disclaimer: The "Vibes Test"

These quantizations have NOT been formally tested for perplexity.

They were compiled as an experiment to see how the model handles shifting bottlenecks. There is no guarantee that they are mathematically optimal or perform flawlessly.

If they pass the vibes test for you, enjoy!

Credits & Acknowledgments


Quantization Recipes

All variants use the same custom tensor buckets: attention, SSM, shared experts, routed experts, embeddings/output, and MTP/NextN tensors.

IQ6_K

Highest quality routed expert quantization in this set.

Layer GroupQuant
Token Embeddings & OutputQ8_0
AttentionQ8_0
SSM Alpha & BetaBF16
SSM OutputQ8_0
Shared ExpertsQ8_0
Routed ExpertsIQ6_K
MTP / NextNQ8_0

IQ5_K

High quality routed expert quantization with IQ5_K experts.

Layer GroupQuant
Token Embeddings & OutputQ8_0
AttentionQ8_0
SSM Alpha & BetaBF16
SSM OutputQ8_0
Shared ExpertsQ8_0
Routed ExpertsIQ5_K
MTP / NextNQ8_0

IQ5_KS

High quality routed expert quantization using IQ5_KS experts.

Layer GroupQuant
Token Embeddings & OutputQ8_0
AttentionQ8_0
SSM Alpha & BetaBF16
SSM OutputQ8_0
Shared ExpertsQ8_0
Routed ExpertsIQ5_KS
MTP / NextNQ8_0

IQ4_K

Balanced 4-bit routed expert quantization with high precision on always-active tensors.

Layer GroupQuant
Token Embeddings & OutputQ8_0
AttentionQ8_0
SSM Alpha & BetaQ8_0
SSM OutputQ8_0
Shared ExpertsQ8_0
Routed ExpertsIQ4_K
MTP / NextNQ8_0

IQ4_KS

Smaller 4-bit routed expert quantization with compressed embeddings, output, and MTP tensors.

Layer GroupQuant
Token Embeddings & OutputIQ6_K
AttentionQ8_0
SSM Alpha & BetaQ8_0
SSM OutputQ8_0
Shared ExpertsQ8_0
Routed ExpertsIQ4_KS
MTP / NextNIQ6_K

IQ4_KSS

Ubergarm-style split routed expert recipe.

Layer GroupQuant
Token Embeddings & OutputIQ6_K
AttentionQ8_0
SSM Alpha & BetaQ8_0
SSM OutputQ8_0
Shared ExpertsQ8_0
Routed Experts DownIQ4_KS
Routed Experts Gate/UpIQ4_KSS
MTP / NextNIQ6_K

IQ3_K

Lower size recipe with IQ3_K routed experts and IQ6_K on many always-active tensors.

Layer GroupQuant
Token Embeddings & OutputIQ6_K
AttentionIQ6_K
SSM Alpha & BetaQ8_0
SSM OutputIQ6_K
Shared ExpertsIQ6_K
Routed ExpertsIQ3_K
MTP / NextNIQ6_K

IQ2_KL

Maximum compression variant in this set.

Layer GroupQuant
Token EmbeddingsIQ4_K
OutputIQ6_K
AttentionIQ6_K
SSM Alpha & BetaIQ6_K
SSM OutputIQ6_K
Shared ExpertsIQ6_K
Routed Experts DownIQ3_KS
Routed Experts Gate/UpIQ2_KL
MTP / NextNIQ6_K

How to Run

  1. 1.Clone and build the ik_llama.cpp fork from ikawrakow/ik_llama.cpp.
  2. 2.Use the compiled llama-server or llama-cli from that specific build.
  3. 3.For chat templating, use the model's embedded template or the community template credited above, depending on your frontend.

Example `llama-server` launch command:

bash
./llama-server -m Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated-IQ4_KS.gguf -c 8192 -ngl 99 -fa --jinja