CoolFace
Modelpublic

Koshkasa/MuXodious_GLM-4.7-Flash-absolute-heresy-mixed-trellis-GGUF

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes231downloads
Model Card

MODEL CARD INCOMPLETE

BENCHMARKS PENDING

!!! quantized for use with ik_llama.cpp and its derivatives !!!

!!! incompatible with mainline llama.cpp as of commit #34af94c !!!

What's that?

An experimental, English-roleplay-oriented, ikllama.cpp-only [MuXodious/GLM-4.7-Flash-absolute-heresy](https://huggingface.co/MuXodious/GLM-4.7-Flash-absolute-heresy) quantization for users who want an IQ3M-sized file but prefer to spend precision on MLA, routing, shared experts, and output-sensitive tensors.

Rationale

I wanted a quant that would fit my GPU with some context and minimal offload. Simple as. This gguf was made to compete with IQ3M/IQ4XS by compressing sparse expert ffn tensors in SOTA low-bit quant types (IQ3KT/IQ4KT) while protecting the most compression-sensitive, architecture critical tensors. I had concerns with mainline quant method compromises - such as shared experts in 3-bit, and the MLA KV "condensation" weights (attnkvamqa) being quanted lower than attnkb (the layer extracting Keys from the shared latent attention vector). Fearing that overcompressed MLA latent projection would mangle reconstructed attention states, I went for a much higher quantization for these. The recipe provided is, however, an exploratory MLA/MoE allocation, not gospel. The proportion of ffnexps parameters in the entire model is 92.46%. And 2.12% for lmhead and embeddings. Meaning EVERYTHING else - the shared exps, the attention tensors - is 3 GB in bf16. These also happen to be quantization sensitive tensors. As a prime example, keeping expert routing weights in bf16 across **the entire gguf** has cost... 7.5 MB over IQ3S.

The imatrix was generated on wrapped natural language english text from eaddario/imatrix-calibration, using kld-sweep-dataset by cmhamiche. The imatrix was not calibrated for STEM, mathematics, code, or non-English languages. I was building it for my use purposes. However, I'm not opposed to making a trellis quant for other use cases if anyone needs it.

UPD: since the IQ3_KT ffn_down_exps also seems to be perfectly functional, it's provided as well.

Mixed Trellis

Component / RoleTensor (Layer)Dense Block 0 (1x)MoE/MLA Blocks (46x)
Global Layerstoken_embd.weight <br> output.weightiq5_k <br> —— <br> iq6_k
MLA Compressedattn_q_a.weight <br> attn_kv_a_mqa.weightbf16 <br> bf16bf16 <br> bf16
MLA Attentionattn_q_b <br> attn_k_b / attn_v_b <br> attn_output.weightq8_0 <br> q8_0 <br> q8_0q8_0 <br> q8_0 <br> iq6_k
MoE Routingffn_gate_inp.weight <br> exp_probs_b— <br> —bf16 <br> f32
MoE Expertsffn_down_exps.weight <br> ffn_gate_exps.weight <br> ffn_up_exps.weight— <br> — <br> —iq4_kt <br> iq3_kt <br> iq3_kt
Shared Expertsffn_down_shexp <br> ffn_gate_shexp <br> ffn_up_shexp— <br> — <br> —q8_0 <br> q8_0 <br> q8_0
Dense MLPffn_down / ffn_gate / ffn_upq8_0—

IQ3_KT

Component / RoleTensor (Layer)Dense Block 0 (1x)MoE/MLA Blocks (46x)
Global Layerstoken_embd.weight <br> output.weightiq5_k <br> —— <br> iq6_k
MLA Compressedattn_q_a.weight <br> attn_kv_a_mqa.weightbf16 <br> bf16bf16 <br> bf16
MLA Attentionattn_q_b <br> attn_k_b / attn_v_b <br> attn_output.weightq8_0 <br> q8_0 <br> q8_0q8_0 <br> q8_0 <br> iq6_k
MoE Routingffn_gate_inp.weight <br> exp_probs_b— <br> —bf16 <br> f32
MoE Expertsffn_down_exps.weight <br> ffn_gate_exps.weight <br> ffn_up_exps.weight— <br> — <br> —iq3_kt <br> iq3_kt <br> iq3_kt
Shared Expertsffn_down_shexp <br> ffn_gate_shexp <br> ffn_up_shexp— <br> — <br> —q8_0 <br> q8_0 <br> q8_0

Cheers

[Z.ai](https://huggingface.co/zai-org) - the base model. [ikawrakow and contributors of ik_llama.cpp](https://github.com/ikawrakow/ik_llama.cpp) - I probably misused your creation. [MuXodious](https://huggingface.co/MuXodious) - for letting the model swear. [cmhamiche](https://github.com/cmhamiche) - for accessible, ready-to-use dataset construction tool. [eaddario](https://huggingface.co/eaddario) - for the imatrix dataset.