CoolFace
Modelpublic

wrldsuksgo2mars/GLM-5.2-EXL3-K3-calibrated-v1

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes70downloads
Model Card

GLM-5.2 EXL3 K3 calibrated v1

This repository contains a calibrated EXL3 K3 variant of `zai-org/GLM-5.2` for GLMRT. The three dense decoder layers, attention, routers, shared experts, embeddings, head, and native MTP layer retain their source tensors. Only the 256 routed experts in base decoder layers 3 through 77 are replaced by EXL3 K3 MCG trellis/suh/svh/mcg tensors.

The artifact uses a single resident expert representation on each of four Spark TP ranks. Every rank retains its 512-wide intermediate slice of every expert; the runtime does not keep native and EXL3 copies resident together.

Quantization

  • —Source: zai-org/GLM-5.2, revision b4734de4facf877f85769a911abafc5283eab3d9
  • —Format: EXL3, 3 physical trellis bits per routed-expert weight
  • —Codebook: MCG
  • —Output-scale search: automatic, folded into the stored rotations
  • —Calibration: 1,080,625 GLM-5.2 tokenizer tokens spanning general text, English and Chinese, code/agentic work, mathematics/reasoning termination, and structured output
  • —Sparse coverage: natural top-8 routes, deterministic adjacent-router recovery for deficient experts, then an explicitly recorded isotropic Hessian residual only for any remaining deficit
  • —Quantizer: the content-pinned GLMRT GPTQModel fork

The calibration, held-out, and screening splits are source-disjoint. The complete immutable plan, projection evidence, retained-native proof, and artifact manifest are maintained by the GLMRT project rather than shipped as large private recovery files in this standard model repository.

Runtime

The artifact is intended for GLMRT's generated SparkInfer SM121 EXL3 K3 TP4 path. Generic Transformers metadata is included, but compatibility with other EXL3 runtimes has not been claimed.

The qualification and performance measurements below were made with the GLMRT v6 inference engine.

Qualification

The exact artifact passed the content-bound GLMRT qualification gate. Both arms used identical coordinator and Spark binaries, the balanced profile, dSpark speculation, and a 400 W coordinator power limit.

Structural and quantizer evidence

  • —Routed projections quantized: 57,600
  • —EXL3 tensors: 230,400
  • —Retained native tensors byte-compared: 1,985
  • —EXL3 tensor payload: 254.00 GiB
  • —TP4 resident payload per Spark: 64.00 GiB
  • —Aggregate Hessian-weighted relative projection error: 0.0063075206
  • —End-to-end serving decision: accepted

Decode and acceptance

WorkloadNVFP4EXL3Ratio
Weighted decode27.016 tok/s28.679 tok/s1.062x
Orchid repeat64.309 tok/s71.439 tok/s1.111x
dSpark accepted drafts76.61%72.65%0.948x
  • —Policy: explicit decode-optimized tradeoff; the ordinary 0.950x acceptance and per-cell prefill floors were not used
  • —Selected acceptance floor: 0.940x of NVFP4
  • —Candidate semantic contracts: 35/35 passed

Prefill

ContextPrefill rowsNVFP4 tok/sEXL3 tok/sRatio
01,024590.0700.71.188x
02,048977.21,031.11.055x
04,0961,488.81,345.80.904x
08,1921,761.91,607.00.912x
016,3841,803.61,683.60.933x
032,7681,773.31,709.70.964x
32,7681,024850.3676.70.796x
32,7682,048965.01,020.31.057x
32,7684,0961,483.61,374.90.927x
32,7688,1921,717.51,645.80.958x
32,76816,3841,714.21,692.80.988x
32,76832,7681,684.11,684.91.000x
65,5361,024768.2663.70.864x
65,5362,048960.31,015.01.057x
65,5364,0961,480.11,349.30.912x
65,5368,1921,646.81,619.90.984x
65,53616,3841,597.31,649.11.032x
65,53632,7681,571.91,622.41.032x
131,0721,024827.1658.40.796x
131,0722,048956.81,010.01.056x
131,0724,0961,362.01,298.30.953x
131,0728,1921,414.91,514.41.070x
131,07216,3841,518.81,523.21.003x
131,07232,7681,484.31,498.61.010x
262,1441,024793.3808.91.020x
262,1442,048904.0969.01.072x
262,1444,0961,221.21,179.60.966x
262,1448,1921,317.51,310.80.995x
262,14416,3841,308.41,324.41.012x
262,14432,7681,290.81,306.81.012x

Minimum prefill cell ratio: 0.796x. Selected per-cell prefill floor: 0.790x of NVFP4.

Tool use and startup

  • —Tool-call score: 125/138 points (NVFP4: 120/138; 1.042x)
  • —Expert resident preload: 17,212.3 ms (NVFP4: 24,551.2 ms; 0.701x)
  • —Full expert service handoff: 17,712.4 ms (NVFP4: 24,975.8 ms; 0.709x)
  • —Native EXL3 parity: TP ranks 0, 1, 2, 3; calibrated layer 3; rows 1, 3, 9, 10, 129, 257, 513, 1025, 2049, 2064

Reproducibility identity

  • —GLMRT WIP engine: wip-exl3-fp32-priority-e9d90fb41fcf-9f564406a95d
  • —SparkInfer revision: f78061114b5575a2969e35b248ed6397424a306c
  • —Coordinator slot SHA-256: e9d90fb41fcfb73b5c90991549458c020f38b7acb8c7dc717edfbc53c1cb86f0
  • —Spark slot SHA-256: 9f564406a95db11707ec9dc7ddfe83079aa0b31be4845a84f2cd7b6b7dac84be
  • —Quantizer execution upgrade SHA-256: 9a1843a7cc13977a21b53ceb9f57bdf938df826d3b4347507946aa186e39f5e1
  • —Qualification evidence SHA-256: 61dc821627b22cb9bc469c8db54bf5d5256787ac46296efda092f71172161c74
  • —Hub revision: assigned and verified immediately after the initial upload

Before publication this marker is replaced mechanically from the signed final structural, quantizer, and serving reports. The publication builder rejects a card that retains the marker or does not cite the exact serving-report hash.

License and attribution

This derivative follows the source model's MIT license. See LICENSE and the original `GLM-5.2` model card for upstream architecture, intended-use, and limitation details.