Deviad/GLM-5.2-shortgpt-pruned-IQ2S-experts-IQ4NL-rest
GLM-5.2 — ShortGPT-pruned, Mixed-Precision GGUF (IQ2S experts · IQ4NL rest)
This is a ShortGPT-pruned re-release of the GLM-5.2 mixed-precision GGUF.
Starting from zai-org/GLM-5.2 (256×22B Mixture-of-Experts, architecture glm-dsa), 12 Transformer blocks were removed by ShortGPT structured layer pruning (block count 79 → 67), and the surviving weights were then re-quantized with llama-quantize's per-tensor mixed-precision workflow using an importance matrix. The MoE expert tensors are stored at IQ2_S (≈2.6 BPW overall) while the dense / attention / norm / embedding / shared-head tensors stay at IQ4_NL.
The goal is the smallest practical memory footprint: the pruned + low-bit model is ≈191 GiB, roughly 18 % smaller than the un-pruned mixed-precision release, while keeping the same exact quantization scheme per tensor.
Model particulars (from GGUF KV metadata)
Pruning
ShortGPT evaluates the importance of each decoder block (via cosine similarity of inputs/outputs) and drops the lowest-importance blocks. On GLM-5.2 this removed 12 blocks (79 → 67), reducing both parameter count and activation memory. Layer indices are sparse afterwards — the retained blocks keep their original indices rather than being renumbered, so the file reports block_count=67.
Quantization mapping
Per-tensor type assignment passed to llama_quantize (same scheme as the sibling un-pruned release):
- Source GGUF:
unsloth/GLM-5.2-GGUFIQ4_NL variant, pruned and re-quantized withallow-requantize+keep-split. - Importance matrix:
imatrix_unsloth.gguf(sourced from Unsloth). - Final size: ≈191 GiB across 9 shards, ≈2.6 BPW.
Files
Filenames include the IQ2_S / IQ4_NL quant tokens so Hugging Face's quantization-variant scanner recognizes the shards (a single quant label is not possible for a mixed-precision quant; both constituent quants are listed).
Usage
Load with any recent llama.cpp build (and compatible runners — LM Studio, Ollama, koboldcpp, etc.) that supports the glm-dsa architecture, MLA attention and IQ2S / IQ4NL dequantization (GPU offload strongly recommended).
llama-server \
-m GLM-5.2-shortgpt-pruned-IQ2_S-IQ4_NL-00001-of-00009.gguf \
--host 0.0.0.0 --port 8080 \
-ngl 999 -c 8192The first shard is the entry point;llama.cppfollows the split-file links to load all 9 shards automatically. Point-mat00001-of-00009.
Provenance
- Base model: zai-org/GLM-5.2 — MIT.
- Source GGUF quantization: Unsloth (
general.quantized_by = Unsloth,general.repo_url = https://huggingface.co/unsloth). - ShortGPT pruning + mixed-precision re-quant with imatrix: Deviad (2026-06-21), on Apple M3 Ultra (Metal build of
llama.cpp).
Disclaimer
This is an aggressive low-bit quantization of an already-pruned model, intended to fit a very large MoE into constrained memory. Expect measurable quality degradation versus the source, both from ShortGPT layer removal and from the IQ2_S expert tensors. Validate on your own tasks before relying on it.
