Koshkasa/MuXodious_GLM-4.7-Flash-absolute-heresy-mixed-trellis-GGUF
MODEL CARD INCOMPLETE
BENCHMARKS PENDING
!!! quantized for use with ik_llama.cpp and its derivatives !!!
!!! incompatible with mainline llama.cpp as of commit #34af94c !!!
What's that?
An experimental, English-roleplay-oriented, ikllama.cpp-only [MuXodious/GLM-4.7-Flash-absolute-heresy](https://huggingface.co/MuXodious/GLM-4.7-Flash-absolute-heresy) quantization for users who want an IQ3M-sized file but prefer to spend precision on MLA, routing, shared experts, and output-sensitive tensors.
Rationale
I wanted a quant that would fit my GPU with some context and minimal offload. Simple as. This gguf was made to compete with IQ3M/IQ4XS by compressing sparse expert ffn tensors in SOTA low-bit quant types (IQ3KT/IQ4KT) while protecting the most compression-sensitive, architecture critical tensors. I had concerns with mainline quant method compromises - such as shared experts in 3-bit, and the MLA KV "condensation" weights (attnkvamqa) being quanted lower than attnkb (the layer extracting Keys from the shared latent attention vector). Fearing that overcompressed MLA latent projection would mangle reconstructed attention states, I went for a much higher quantization for these. The recipe provided is, however, an exploratory MLA/MoE allocation, not gospel. The proportion of ffnexps parameters in the entire model is 92.46%. And 2.12% for lmhead and embeddings. Meaning EVERYTHING else - the shared exps, the attention tensors - is 3 GB in bf16. These also happen to be quantization sensitive tensors. As a prime example, keeping expert routing weights in bf16 across **the entire gguf** has cost... 7.5 MB over IQ3S.
The imatrix was generated on wrapped natural language english text from eaddario/imatrix-calibration, using kld-sweep-dataset by cmhamiche. The imatrix was not calibrated for STEM, mathematics, code, or non-English languages. I was building it for my use purposes. However, I'm not opposed to making a trellis quant for other use cases if anyone needs it.
UPD: since the IQ3_KT ffn_down_exps also seems to be perfectly functional, it's provided as well.
Mixed Trellis
IQ3_KT
Cheers
[Z.ai](https://huggingface.co/zai-org) - the base model. [ikawrakow and contributors of ik_llama.cpp](https://github.com/ikawrakow/ik_llama.cpp) - I probably misused your creation. [MuXodious](https://huggingface.co/MuXodious) - for letting the model swear. [cmhamiche](https://github.com/cmhamiche) - for accessible, ready-to-use dataset construction tool. [eaddario](https://huggingface.co/eaddario) - for the imatrix dataset.
