CoolFace
Modelpublic

inference-optimization/GLM-5.3-Flash-MEP50

sourceHugging Faceupdated 1mo agoView on Hugging Face
4likes165downloads
Model Card

GLM-5.3-Flash - 50% Expert Pruned

50% of the MoE experts pruned by router-weight magnitude using compressed-tensors.

  • —Base model: zai-org/GLM-5.3-Flash
  • —Sparsity: 50% of routed experts removed per layer
  • —Layers pruned: all 43 MoE layers (body layers 3-44)
  • —MTP layer: retained exactly as-is
  • —Shared experts: untouched
  • —Vision tower: untouched (dense ViT MLP, no MoE experts)

Reproduction