CoolFace
Modelpublic

brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78

sourceHugging Facemitupdated 1mo agoView on Hugging Face
9likes3.3kdownloads
Model Card

GLM-5.2 EXL3 TR3 v4 — 3.5 bpw mixed experts, MTP-78

Current serving status — read this first

There is no universal current default launcher across the historical profiles retained in this card. For a new deployment, do not infer a default from an older section:

ProfileStatusScope
Portable r17 graph/NVFP4 packageRelease candidate; not yet the defaultReproducible build and focused source/ABI tests; end-to-end GPU replay is still required
Dedicated-R7 r17 candidate used for the results belowLocally qualified evidence sourceTP4/DCP4/MTP3, full graphs, NVFP4 DS-MLA, tasks and 260K retrieval
r17 exact-G64-Q FP8 readerProvisional quality experimentEager KLD result; not the graph/speed-qualified serving profile
Gilded Gnosis r34Historical qualified reference (2026-08-10)TP4/DCP1, 65,536-token limit; not the present DCP4 long-context candidate
Earlier max-KV launchersHistorical measurementsPreserve provenance only; use their exact image and configuration when reproducing them

The portable r17 package becomes the recommended downloadable path only after it repeats the end-to-end graph, task, 260K retrieval, and numerical gates on the corrected dual-ABI build.

Sealed r17 FP8-reader KLD comparison (2026-08-22)

An immutable five-fresh-boot acceptance panel plus a factorial G64-Q-only follow-up compared the r17 FP8 MLA reader paths. Lower KLD is better.

reader regimefive-run mean KLDmediansample SDresult vs stock
stock reader flags0.060333840.060802580.00122809control
exact G64 Q only0.059730210.058940770.00233983-1.00% KLD; `<0.060` gate passed
BF16 P.V0.061553030.060901190.00197113+2.02% KLD
exact G64 Q + BF16 P.V0.061058020.060361780.00137015+1.20% KLD

The five G64-Q-only values were 0.05855139, 0.06383442, 0.05894077, 0.05927899, and 0.05804549. The arithmetic mean passed the predeclared <0.060 gate, but one high boot increased sample SD and the five-boot Welch contrast with stock remains inconclusive (p=0.6276).

Provisional quality launcher: server.sh launches the digest-pinned r17 runtime with exact G64 Q-only enabled, MTP-3, eager execution, and a fixed 4 GiB KV reservation per rank. The validated startup initialized 313,856 FP8 MLA KV tokens at a 262,144-token model limit (1.20x maximum-request concurrency). This is not the speed-qualified release default. One attempted routing entered the ordinary per-expert branch and hit nonzero during capture; that does not show that r17 or its EXL3 fallback cannot run in CUDA graphs. Full r17 graph replay with nvfp4_ds_mla and measured speed are mandatory before the quality battery is rerun. The three required source overlays and their checksums ship in `runtime/r17-g64-q-only`. Neither BF16 P.V combination advances. G64 uses the existing Q-row padding for eight exact G64 scales, leaves BF16 P.V and BF16 QK disabled, and adds no KV-record bytes, ABI change, or dynamic-shared-memory growth.

The panel used one pinned 2,048-token WikiText window, 2,047 full-vocabulary teacher-forced positions, the official BF16-reference logits, FP8 MLA KV with BF16 RoPE, TP4/DCP4, and 300 W per GPU (1,200 W aggregate). Boot repeats on one window are not independent prompts, so the report does not claim general task quality or production qualification.

Graph-qualified r17 NVFP4 MLA profile (2026-08-22)

This checkpoint has now been qualified on an r17-derived, dedicated R7 runtime with full CUDA graphs. This section is separate from the FP8-reader KLD panel above: the KLD values, KV formats, and runtime paths are not interchangeable.

Validated launch contract

  • —4x RTX PRO 6000 Blackwell, TP4 / DCP4 / MTP3
  • —dedicated projection-exact exl3_moe_r7_fused
  • —checkpoint-native independent gate/up/down K3/K4/K5 assignments
  • —sorted router probabilities preserved in FP32
  • —nvfp4_ds_mla KV using nvfp4_mla_outer_scales.json
  • —BF16 RoPE (KV_FP8_ROPE=0) and BF16x3 router GEMM
  • —ONLINE_QUANT=none for the tested checkpoint layout
  • —enforce_eager=False
  • —CUDAGraphMode.FULL_AND_PIECEWISE; FULL graphs captured on all ranks
  • —4 GiB KV reservation per rank: 463,872 physical KV tokens
  • —262,144-token model limit

Use the files in `runtime/r17-graph-nvfp4` for this profile. The self-contained Dockerfile starts from the pinned public r17 digest and applies the reviewed vLLM and ExLlamaV3 patch series. The wrapper checks the calibrated-scale hash, selected launch flags, graph log markers, KV allocation, and graph replay before declaring the server qualified; it does not independently attest every source or checkpoint hash.

Publication gate: the portable image passes its build, ABI assertions, and 68 focused vLLM tests, but has not yet repeated the end-to-end GPU battery. The measurements below belong to the earlier validated candidate whose effective overlays this portable build reconstructs; do not promote the portable image to the default until replayed.

Measured performance

Four nominal 300 W GPU caps, 1,200 W aggregate:

ProfileResult
Prefill 8K1,203 tok/s
Prefill 64K1,161 tok/s
Prefill 128K1,107 tok/s
Decode C174.0 tok/s
Decode C2118.7 tok/s
Decode C4192.5 tok/s

Five-run task validation

Post-reseat quality used nominal caps of 300/300/300/275 W:

ProfileVerdictMean output tokensAggregate generation
LAVD5/5 correct (2 exact, 3 near)8,728.859.27 tok/s
Estonia5/5 pass1,744.255.65 tok/s
Hotel5/5 exact20,143.855.53 tok/s
Estonia-long5/5 pass1,951.454.82 tok/s

No successful run hit its output-token limit. The small TTFT figures in the task receipt are prefix-cache affected and are not uncached-prefill claims.

260K retrieval

Five 260,000-token prompts retrieved their planted record exactly at 5%, 25%, 50%, 75%, and 95% depth. The later probes benefited from an increasing shared prefix, so their elapsed times are not an uncached-prefill comparison.

BF16-reference KLD

Four comparable eager-harness 2,047-position observations exist for the exact dedicated-R7/FP32-route/static-scale/NVFP4 configuration: 0.0639764262, 0.0746190367, 0.0704009716, and 0.0676657456. Their descriptive mean is 0.0691655450. These comprise one sentinel plus three completed runs from a planned five-run panel. Run 4's log ends during checkpoint loading without a recorded cause, and run 5 never began, so there is no five-run acceptance mean. The numerical harness used eager execution with CUDA graphs disabled; graph qualification comes from the separate serving/task/needle receipts.

This numerical result should not be conflated with the five-run FP8 MLA reader result above (0.05973021 for exact G64-Q-only). The FP8 reader study used a different KV format and was explicitly not the graph/speed-qualified serving default.

Hardware and claim boundaries

The successful post-reseat batteries recorded no Xid, PCIe/AER, CUDA, traceback, engine-fatal, hardware-slowdown, thermal-slowdown, or power-brake event. Maximum observed temperature was 89 C; GPU 3/C1 reached 76 C. Telemetry can report short samples above a configured power limit, so the stated caps are nominal configuration limits rather than hard instantaneous clamps.

This qualifies the exact local checkpoint file set identified by the published configuration, quantization-config, safetensors-index, and scale hashes; the tested image; TP4/DCP4/MTP3 topology; NVFP4 DS-MLA KV contract; maximum four sequences; and contexts through 260,000 tokens on the reference host. It does not bind that tested file set to a particular Hub commit or establish multi-prompt KLD generalization, concurrency above four, FP8 KV graph qualification, or other GPU topologies.

Full receipts:

The vLLM integration is tracked in local-inference-lab/vllm PR #300. The dedicated kernel companion is tracked in turboderp-org/exllamav3 PR #246, which still requires a port from its v0.0.43 base to current upstream.

Checkpoint layout update (2026-08-10): BF16 shared experts, merged online K6

The shared-expert MLP tensors are now stored as the original BF16 weights instead of pre-encoded EXL3 payloads, and the serving stack encodes them as one merged K6 payload at load. Routed R7 expert payloads and rotations are byte-identical to the previous revision; nothing else about the quantization changed.

If you pulled this repo before 2026-08-10, re-download model.safetensors.index.json, config.json, quantization_config.json, the model-layer-003 .. model-layer-078 carrier shards, and the new model-sharedbf16.safetensors. hf download picks up the delta on its own.

What changed and why

Festr isolated a decode-throughput loss to the previous shared-expert layout. mlp.shared_experts.{gate_proj,up_proj,down_proj} were stored as three separately encoded K6 EXL3 trellis payloads per MoE layer (layers 3-77; the layer-78 MTP draft shared experts were already BF16). Shared gate and up as two separate payloads force two small-M GEMM launches per routed layer per decode step where a merged payload needs one -- 75 extra kernel launches per decode step at MTP0. His diagnostic, TP4 / DCP1 / MTP0, split vs merged shared gate+up:

decode tok/sC1C4C8
split gate/up (previous layout)50.62157.34258.81
merged gate+up53.86169.07281.72

This revision therefore stores the shared experts unencoded:

  • —A new shard model-sharedbf16.safetensors (~5.7 GB) carries the 228 shared-expert BF16 tensors: gate/up/down for layers 3-78, including the MTP-78 draft layer.
  • —The 76 carrier shards model-layer-003 .. model-layer-078 are rewritten without their shared-expert entries. All other shards are unchanged.
  • —tensor_storage in quantization_config.json (and the copy embedded in config.json) drops its 225 shared-expert module entries, so a loader sees the shared experts as ordinary BF16 modules.
  • —Routed experts (r7-experts-layer-*.safetensors) and all rotations: byte-identical, untouched.

Serving

Set ONLINE_QUANT=exl3-b6. At load the runtime concatenates gate and up while still BF16 and encodes one merged K6 payload per layer; the encoded result lands in the JIT/weight cache, so the cost is paid once, on first load. One payload, one launch.

Without ONLINE_QUANT the checkpoint still serves, with the shared experts running in plain BF16: correct output, roughly 3.5 GB more weight memory model-wide than the online-encoded path (5.74 GB BF16 versus 2.20 GB encoded), and none of the merged-launch decode gain.

Quality: single pass from source

The BF16 tensors are the original shared-expert weights -- verified bit-identical across the willfalco 3.42-bpw checkpoint and willfalco 3.25-bpw checkpoint lineage checkpoints by independent ranged-read sha256 sampling. The online merged K6 encode is one quantization pass from that source, exactly as the previous layout's offline K6 encode was. Nothing is re-quantized from an already-quantized representation.

KLD, 5-run gate against the same BF16 reference logits:

layoutmeansd
previous (offline split K6 shared)0.0624500.001533
this (BF16 shared, ONLINE_QUANT=exl3-b6)0.0642500.000383

Verdict: PASS with documented delta -- 0.064250 +/- 0.000383 (5 runs, full exl3-b6 serving policy) vs 0.062450 +/- 0.001533 for the previous layout (5-run gate). Delta +0.0018; the new mean sits below the previous layout's own worst run (0.064910). Traded for +9.5% mean decode throughput and +20.5% KV-cache capacity.

Decode at MTP-3 on the reference rig (4x RTX PRO 6000, TP4 + DCP4):

configprevious layout (tok/s)this layout (tok/s)delta
C1 decode, MTP-3 (ctx 0-32k avg)61.3 tok/s69.8 tok/s+13.9%
C2 decode, MTP-3 (ctx 0-32k avg)90.7 tok/s99.1 tok/s+9.3%
C4 decode, MTP-3 (ctx 0-32k avg)127.2 tok/s136.8 tok/s+7.5%
KV cache @ 262,144 MML (nvfp4dsmla)432,384 tokens (1.65x)520,960 tokens (1.99x)+20.5%
Prefill 8k-32k1309-1365 tok/s1301-1345 tok/s-0.6 to -1.5% (within run variance)

Historical qualified reference (2026-08-10): Gilded Gnosis r34

On 2026-08-10 — the same day this layout shipped — the local-inference-lab release pipeline published Gilded Gnosis r34 with this checkpoint, at this exact revision, as its qualified reference:

text
voipmonitor/vllm:gilded-gnosis-v20-vllm4d006a4-b12xcd3ce19-fi1ac6942-cu132-20260810-r34
sha256:820181fbbc975cd5291c411cda9771d58fecee1636d916f508f47230df20592b
qualified: brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78@9ab9579774cc432df91567a36f6e9e863e0d4c9f

The r34 runtime keeps the routed experts in their serialized K3/K4/K5 Trellis formats and encodes the BF16 shared experts into cached merged K6 projections (this repo's layout, consumed as intended). Loader contract: InstantTensor BUFFERED with borrowed-buffer consumption. Qualified profile: TP4/DCP1, B12X A16, B12X sparse MLA, NVFP4 DS-MLA KV, MTP-3, 8 sequences, graph cap 32, model limit 65,536, GMU 0.98.

Release-gate measurements (their receipt, not mine):

profileC1C4C8prefill 8KKV tokens
MTP0 / GMU .9753.80171.30283.053,25382,816
MTP3 / GMU .98121.25297.69436.233,23975,072

MTP-3 strict acceptance 65.44%. FULL decode graphs covered every configured size; target verification, all three MTP forwards, and prefill remained graph-captured. Focused vLLM, B12X host/GPU, runtime-contract, startup, deterministic-output, and checksum gates passed.

  • —Runbook and immutable Compose: <https://github.com/local-inference-lab/rtx6kpro/blob/master/models/glm5.2_v20.md>
  • —Source merge contract and machine-readable evidence: <https://github.com/local-inference-lab/rtx6kpro/issues/33>

Scope note, in the release's own words: the r34 receipt does not qualify R7 at DCP>1, standard NVFP4, or NF3 performance. The long-context DCP4 profile documented above (262,144-token context, 520,960 KV tokens) is the author's own measured configuration, validated by the KLD and throughput gates in this card, not by the r34 receipt.

Credit where it belongs: the split-payload decode loss was isolated by Festr, whose analysis produced this layout, and the qualification is the work of the local-inference-lab Discord community and its release engineering. Same-day pipeline from proposal to shipped checkpoint to qualified release — that is what a receipts-first community looks like.

Previous layout

The pre-update revision remains available at `c55c1cd4ca42` if you need the old offline-encoded shared payloads.

The exact source lineage, independent TP4 validation, conversion tool, and focused tests are recorded in BF16_SHARED_ONLINE_K6_VALIDATION.md.

Historical and provisional launchers

Provisional eager FP8/G64 research image (digest-pinned r17 base):

voipmonitor/vllm@sha256:c5e96c5bcc5a073f7ce6b56173d88538de3a416900cff97c88b4bf7967fe1dc0
bash
docker pull voipmonitor/vllm@sha256:c5e96c5bcc5a073f7ce6b56173d88538de3a416900cff97c88b4bf7967fe1dc0
hf download brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78 --local-dir ./glm52-tr3v4
cd ./glm52-tr3v4 && ./server.sh start
FileWhat
runtime/r17-g64-q-onlyExact reader overlays and checksums
SERVING.mdImage, serve parameters, gotchas
RESULTS.mdMeasured KLD and throughput
docker-compose.ymlHistorical pre-G64 reference deployment
server.shProvisional r17 G64-Q-only start / stop / logs wrapper

Reference rig: 4x RTX PRO 6000 Blackwell (SM120a, 96 GB, PCIe), TP4 + DCP4, MTP-3. KLD 0.061282 +/- 0.001376 (fp8 KV, BF16 RoPE, 5 runs). Decode 60.5 / 81.5 / 115.5 tok/s at c1/c2/c4; prefill 1595 tok/s at 8K.

The image is built for sm_120a (Blackwell). It will not run elsewhere without a rebuild, and the checkpoint needs a mixed-bit loader -- see the loader compatibility warning below.

Fused MoE path (long context)

An alternative serving path that runs the routed experts on SparkInfer's fused mixed-Trellis MoE kernel. It is opt-in (VLLM_EXL3_R7_FUSED=1) and targets long-context work: more KV capacity per GB than the numbers in "Run it" above, at a slightly higher KLD because it pairs with the 4-bit nvfp4_ds_mla cache rather than fp8.

Measured on 4x RTX PRO 6000 Blackwell (SM120a, 96 GB, PCIe Gen5, no NVLink), TP4 + DCP4, MTP-3:

value
KV cache355,328 tokens
Max concurrency1.36x at 262,144 tokens/request
KLD vs BF16 reference logits0.069527 (4-run mean; fifth interrupted)
Prefill @ 8K1,725 tok/s
Decode75.6 tok/s

Full configuration for every number above: max_model_len 262144, gpu_memory_utilization 0.955, max_num_batched_tokens 2048, max_num_seqs 4, CUDA-graph size 32, kv_cache_dtype=nvfp4_ds_mla with outer scales, BF16 RoPE, VLLM_EXL3_PREFILL_BLOCK_M=64, VLLM_EXL3_PREFILL_CAPACITY=1024, VLLM_DCP_INDEXER_SHARDS=4, 48 fused layers. Prefill measured with a unique prompt prefix so the prefix cache cannot serve it; decode with ignore_eos over 256 tokens.

Non-fused reference on the identical rig and settings: 497,408 tokens KV, 1,757 tok/s prefill, 73.5 tok/s decode. The fused path trades KV capacity for decode throughput.

The 0.061282 figure in "Run it" is fp8 KV cache; the 0.069527 here is nvfp4 KV. They are different cache formats measured against the same reference logits, so compare them with that in mind rather than as a regression.

The outer scales file is required

`nvfp4_mla_outer_scales.json` now ships in this repo. Mount it and point VLLM_NVFP4_MLA_SCALES_FILE at it. Measured on the identical build, changing nothing else:

nvfp4 KVwith outer scaleswithout
BF16 RoPE0.0695270.099717
FP8 RoPE0.0755420.114359

Omitting it costs about 30% KLD for no memory or throughput benefit. It is a per-layer outer-scale calibration (wikitext-2-raw-v1, 2,048 context).

The table also shows why BF16 RoPE is the default here: FP8 RoPE yields roughly 16% more KV tokens but measured +8.7% KLD with scales applied.

What had to be fixed

The fused path was previously unusable on this checkpoint and would have produced garbage output, not a mild regression. SparkInfer bounded the FC2 (down-projection) tier-local expert index using the FC1 slot count. Because this checkpoint chooses the trellis bit width per (expert, projection), FC2 holds more experts than FC1 -- tier1 carries 231 down-projection experts against 77 gate/up -- so most down tiles were rejected and their output silently dropped: 6,653 of 12,288 down-projection expert slots, 54.14%, across the 48 fused layers.

The failure was fluent rather than obviously broken. zero_fc2_output=False and the FC2 buffer aliases rotation_gate, so a rejected tile left gate-rotated hidden states in place, which were then Hadamard-rotated, scaled by down_svh, router-weighted and accumulated. Measured KLD 2.36 with output that hallucinated case law and degenerated into verbatim repetition.

The fix is patches/patch_sparkinfer_projection_tiers.py (fc1_bound_ok).

Files

FileWhat
SERVING_FUSED.mdFull fused-path guide, every knob and its measured justification
nvfp4_mla_outer_scales.jsonRequired outer-scale calibration
serve.shReference serve script for the fused path
docker-compose.yamlReference compose deployment
patches/Patches needed to build your own fused image

Docker image

verdictai/glm52-exl3-sparkinfer:v39-r28-r7fused-broadcast-cu132-sm120a
bash
docker pull verdictai/glm52-exl3-sparkinfer:v39-r28-r7fused-broadcast-cu132-sm120a

Built for sm_120a (Blackwell); it will not run on other architectures without a rebuild. Every patch in patches/ is already applied inside it -- that directory is only needed if you are building your own image from an r28-or-later SparkInfer base.

An r28-or-later base is required, not merely preferred: patch_r7_broadcast_rotations.py depends on SparkInfer ABI-v6 broadcast_suh/broadcast_svh, which earlier bases do not expose. Without it the loader expands one shared rotation row per layer into 256 identical copies, costing about 0.5 GiB per rank.

Caveats

  • —KLD 0.069527 is a 4-run mean; the runner was interrupted before a fifth. The separation from the unscaled 0.099717 is far larger than the run-to-run spread, but treat the third decimal as provisional.
  • —Prefill and decode figures are single probes on a rig that has shown double-digit container-to-container variance. Treat them as indicative.
  • —The occupancy patches (patch_moe_deadscale_2cta.py, patch_moe_stages3_only.py) raise the fused kernel from 8 to 15.4 warps/SM, but measured at parity end-to-end on this rig. They are included for completeness, not as a speed claim.

What is different about this quantization

Full-width down-projection encoding

Each expert down projection is encoded jointly across its full 2048-channel input dimension. It is not independently quantized as four serving-specific 512-channel slices. Error correction can therefore compensate across the whole tensor before a loader slices it for tensor parallelism.

Down calibration uses reconstructed gate/up outputs

Gate and up are encoded at candidate bit widths and reconstructed through the same quantized representation that will be stored. Their reconstructed SwiGLU output supplies the conditional calibration input for down. Down is therefore optimized for the quantized gate/up tensors it will actually follow, not for unquantized BF16 gate/up outputs.

Router-mass-weighted exact bit budgeting

Each layer contains 768 routed-expert tensors. Starting all tensors at 3 bits uses 2,304 bit units; the exact 3.5 bpw target is 2,688 units, leaving exactly 384 one-bit upgrades. Candidate 3-, 4-, and 5-bit losses are weighted by the captured float32 routing mass. A dynamic program spends all 384 upgrades, and a tensor may receive a 4→5 upgrade only after its 3→4 upgrade is selected.

This is why an equal-average “barbell” split of very high and very low bit widths was rejected: quantization error falls smoothly with bit width, so the high end wastes marginal bits while the low end crosses a steep error cliff.

Expert-private intermediate reordering

Five intermediate-channel orderings are considered for each expert. The chosen ordering is baked consistently into gate output, up output, and down input. SwiGLU is elementwise, so a consistent permutation is functionally free at serving time while giving the error-correcting walk a better row order.

Shared residual rotations, private intermediate rotations

The 6144-dimensional residual/model-space sides are shared per layer:

  • —gate_up_suh is shared by all gate/up inputs in the layer.
  • —down_svh is shared by all down outputs in the layer.

The 2048-dimensional intermediate sides remain private to each expert:

  • —gate output (gate_svh)
  • —up output (up_svh)
  • —down input (down_suh)

The residual space must be common because tokens enter all routed experts in one coordinate system and expert outputs are routing-weighted and summed back into that system. The intermediate space exists only inside one expert, so each expert can choose the sign-vector draw that best conditions its own weights and activations without imposing a layer-wide compromise. Twelve candidate draws are searched.

Per-128-channel scales folded into the stored representation

Scales are searched independently on a 128-channel grid and folded into the existing per-element sign representation. This gives finer conditioning without a separate runtime scale tensor.

Topology-neutral schema v2

Routed-expert tensors are stored whole with a per-tensor bit map. Tensor- or expert-parallel slicing happens at load time on 128-channel boundaries; no four-GPU topology is baked into the files.

Authoritative corrected local build

The corrected model did not rerun the expensive bit/permutation search. It recovered and froze all 75 layer decisions from R10, corrected the absolute normalization/global-scaling math, and rebuilt layers 3–77 causally on four local SM120 GPUs. Layer L+1 was calibrated from the corrected installed output of layer L.

For every layer, the successful supervisor performed:

  1. 1.a four-GPU attention/router-only flat capture;
  2. 2.one streamed absolute-normalization/GSS fit;
  3. 3.four pinned GPU workers consuming a dynamic 256-expert queue;
  4. 4.a corrected successor forward to create the next layer's input state;
  5. 5.atomic promotion and a durable layer seal; and
  6. 6.reclamation of reproducible capture, predecessor-state, expert-mini-shard, and no-longer-needed BF16 source-window data.

The complete code, tests, decisions, receipts, and storage runbook are in `reproducibility/local-corrected-v1`. The 75 routed shards total 317,347,848,944 bytes and their sidecar manifests total 350,486,725 bytes. All 75 layer seals are included.

Historical rental-box R10 run

The files under `reproducibility/r10` record the B300 capture/search and first encoding attempt. They remain important provenance for the corpus, deterministic prompt plan, inventories, and frozen allocation decisions, but the old encoder is superseded for reproducing corrected routed payload bytes.

1. Seal inputs and runtime

The BF16 source, carrier checkpoint, numeric core, compiled extension, package versions, and runtime Python files were inventoried before timed work. One scalar SafeTensors serializer defect was repaired after that launch-time seal; the published runtime_inventory.r10.json is retained as the literal launch record, while DEPLOYED_CODE.sha256 binds the exact post-repair code that produced the completed layer shards.

2. Flat capture on NVMe

r10_capture.py performs a BF16 source forward and writes one flat capture per MoE layer through LayerCalibRAM/memory-mapped storage. The 75 captures total about 970 GB, so they live on NVMe rather than /dev/shm.

The exact corpus file is `calibration/reap_recall_calib.jsonl`:

  • —12,228 JSONL records
  • —34,002,059 bytes
  • —SHA-256 cf247acc7c5da9f0600c7d6ab3b7c2fcfc54ec30b794e3b6047559285fa44df4

The deterministic capture plan selected 1,773 prompts totaling 1,049,589 tokens. The exact selection and the complete 75-layer capture summary are in `reproducibility/r10/capture`.

Capture and encoding were operationally pipelined: GPU 0 published captures in layer order while the other GPUs consumed already-sealed captures. After all 75 captures completed, GPU 0 joined the encoder pool.

3. Dynamic 12-worker encoding

Six B300 GPUs run two pinned workers each. Workers pull completed layer captures from a SQLite dynamic queue rather than receiving static layer ranges. The deployed settings are:

text
layers             3-77
workers            12 (two per GPU)
GPUs               6
rotation draws     12
held-out rows      4096
minimum fit rows   1024
row chunk          4096
factor cache       512 MiB per worker
sigma regularizer  0.025
CPU threads        36 per worker

Each completed layer is emitted atomically as one r7-experts-layer-NNN.safetensors shard plus its JSON manifest. A layer does not enter the completed queue state until both artifacts exist.

4. Assemble, upload, and verify

The final assembler hard-links complete expert shards, hard-links clean carrier shards, rewrites only carrier shards that mix retained and replaced tensors, and creates the final tensor index and manifest. Progressive uploads use the same final expert-shard basenames, so the authoritative final folder upload reuses the already-present Hub objects.

What was removed for wall-clock speed

The accuracy design above was not weakened. Work that did not change emitted bytes was removed from the timed encoder: payload/runtime re-hashing, routing audits, fixed-point successor passes, install-and-forward checks, pack/unpack repeat decodes, functional oracles, and assembly/conversion. Search, allocation, reconstruction-based down calibration, rotations, permutations, and final encoding remain.

The result is sealed by construction and written atomically. Full checkpoint assembly and structural verification occur after all 75 layers complete. There is deliberately no claim that expensive functional evaluation ran during quantization.

Historical rental-box performance

The live machine is:

  • —6× NVIDIA B300 SXM6 AC, 275,040 MiB each, compute capability 10.3
  • —Python 3.12.3
  • —PyTorch 2.12.1+cu132
  • —Transformers 5.12.1
  • —CUDA 13.2
  • —exllamav3 numeric extension built from v0.0.43 for sm_103
  • —4.295 TB NVMe filesystem and a 1.509 TB memory cgroup

Uncached completed layers have taken roughly 130–138 minutes per worker, including search, all 256 expert probes, exact allocation, final encoding, and atomic emission. Twelve concurrent slots turn that into waves; it is not a 75× serial runtime. The uploader and memory-cache reclaimer run at reduced priority so GPU encoding remains dominant.

Corrected KLD result

The corrected local checkpoint was measured five times against the same BF16 reference logits, with standard FP8 KV cache, BF16 RoPE, TP4/DCP4, one 2,048-token context, and 2,047 scored positions per run:

text
run 1  0.0616431846
run 2  0.0623442891
run 3  0.0626262729
run 4  0.0604832642
run 5  0.0593142137
mean   0.0612822449
sd     0.0013762398

The reference-logits SHA-256 is 87f992a689c054a0548a4b3863da6c809f9239beacd5786d0401e45904fec063. The exact result JSON, raw run logs, evaluation-code hashes, runner, and scoring script are published with the corrected bundle.

Reproducing the corrected run

Start with the exact guide in `reproducibility/local-corrected-v1/README.md`. It explains the required checkpoint identities, decision recovery, local inventory rebinding, layer-3 causal anchor, rolling execution, seal adoption, NVMe free-space frontier, BF16 shard paging, safe cleanup order, assembly, and KLD command. The old `reproducibility/r10/README.md` is historical lineage, not the corrected byte-production runbook.

Tuning: prefill route block size

The single largest configuration win found on this checkpoint. The default VLLM_EXL3_PREFILL_BLOCK_M=64 leaves substantial performance on the table.

Measured end-to-end on the reference rig (4x RTX PRO 6000, TP4 + DCP4, MTP-3, nvfp4_ds_mla KV, max_model_len 262144, max_num_batched_tokens 2048, util 0.95). Output verified coherent and unchanged at every setting:

`VLLM_EXL3_PREFILL_BLOCK_M`64 (default)3216**8**
Prefill 8K (tok/s)1157140314711674
Prefill 64K (tok/s)1134132013971600
Prefill 128K (tok/s)1112121113081514
Decode (tok/s)57.8459.5563.4470.43
KV capacity (tokens)441,344446,720448,512448,512

At block_m=8: prefill +44.7% / +41.1% / +36.2%, decode +21.8%, KV +1.6%. Prefill, decode and KV all improve together -- there is no trade-off to balance.

Why

Nsight Compute on the isolated MoE kernel (single GPU, synthetic weights, so no TP collectives to deadlock the profiler) at the stock block_m=64:

DRAM Throughput          24.03%      <- memory not saturated
Compute (SM) Throughput  35.96%      <- compute not saturated either
Registers Per Thread        244
Dynamic Shared Memory    101.38 KB of 102.40 KB configured
Block Limit Registers         1
Block Limit Shared Mem        1
Theoretical / Achieved Occupancy   16.67%

Dropping to block_m=16 moves DRAM throughput 24.03% -> 64.22% and halves kernel duration (1530 -> 575 us), with registers 244 -> 144 and shared memory 101.38 -> 52.22 KB.

Occupancy does not change (16.67% either way). The kernel is a cooperative one-grid launch -- grid size equals the SM count, so exactly one block per SM by construction. The gain is instead the zero padding a 64-row route block carries when it holds far fewer live rows; smaller blocks waste less.

block_m below 32 additionally requires a register-count table entry for the cta_m_blocks=1 specialization ((256, 1, 16, 4, False)), absent upstream; without it the launch model raises missing W4A16 register count for NVFP4 BF16 specialization.

Levers that did NOT help

Measured and rejected, so nobody repeats them:

LeverResult
FC1/FC2 tile configtile_n=128 already optimal (tn=64 was 18% slower at m=128; tn=256 unsupported)
_STAGES pipeline depth4 already optimal; 3 within noise; 2 was 27% slower
VLLM_EXL3_TRELLIS_MAX_M 32 -> 64within noise (+0.7-1.0% prefill, decode flat)
max_num_batched_tokens 2048 -> 3072trade-off, not a win: prefill +2.7-8.2% but KV -18% (448,512 -> 366,336)

KV cache

The serving stack uses MLA (kv_lora_rank=512, qk_rope_head_dim=64), so the KV cache is one latent per token per layer rather than per-head K and V, and DCP4 shards it across the four GPUs (--dcp-kv-cache-interleave-size 64).

Supported `--kv-cache-dtype` (the bundled launcher rejects anything else): nvfp4_ds_mla (historical bundled-launcher default, 4-bit MLA latent), fp8_ds_mla, fp8.

Measured capacity

Every row below is a number printed by vLLM at startup on the reference rig (4x RTX PRO 6000, 96 GB, TP4 + DCP4, MTP-3, CUDA graphs on, nvfp4_ds_mla KV). A KV number is meaningless without its utilization, KV dtype, context length and batch budget, so all of them are listed.

utilKV dtypemax_model_lenmax_num_batched_tokensMTPKV capacity
0.95nvfp4dsmla8,1922,0483457,728 tokens
0.95nvfp4dsmlanot recorded3,0723340,224 tokens
0.97nvfp4_ds_mla262,1441,2803502,016 tokens (1.91x)
0.97nvfp4dsmla262,1441,5363489,984 tokens
0.97nvfp4dsmla262,1443,0723417,280 tokens
0.955nvfp4dsmla262,1442,0483262,912 tokens

The first row is a historical dedicated-R7 profile (exl3_moe_r7_fused), measured 2026-08-05. The second is the configuration the throughput numbers in RESULTS.md were taken under; its max_model_len was not recorded alongside the figure, so treat it as indicative rather than reproducible.

Capacity moves inversely with max_num_batched_tokens: dropping 3,072 -> 2,048 returned roughly 117k tokens of KV, because the prefill scratch arena and profiling peak shrink with the batch budget.

About the 1.13M figure

server.sh carries a comment claiming ~1,132,544 tokens (2.16x at 524K context) at GPU_MEMORY_UTILIZATION=0.96. Do not read that as usable serving capacity. It is an auto-profile ceiling measured without the speculative decode path engaged, and a configuration that reaches it is not one you would serve from. It is retained in the script's comment for provenance only and is not reproduced as a headline number here.

Pinning a smaller cache

Capacity is auto-profiled at startup. Leave NUM_GPU_BLOCKS_OVERRIDE empty to take the maximum the utilization allows, or set a positive integer to pin a smaller cache and leave memory for other work on the same GPUs.

Note the KLD figure in RESULTS.md was measured with fp8 KV and BF16 RoPE, while throughput and the capacities above use `nvfp4_ds_mla`. They are not interchangeable.

Max-KV serving profile (2026-08-20)

A tuned profile that reaches 502,016 KV tokens at 262,144 max-model-len — a +41% larger cache than the previously published 355,328 — while keeping MTP-3 on DCP4 and full CUDA graphs. Full sweep, traps, and per-lever attribution are in KV_SPEED_RESULTS.md. Turn-key launch: serve.sh and docker-compose.yaml.

GPU_MEMORY_UTILIZATION=0.97 MAX_BATCHED_TOKENS=1280 MAX_NUM_SEQS=4 GRAPH=16

metricthis profilepreviously published
GPU KV cache @ 262,144502,016 tokens (1.91x)355,328 (1.36x)
Prefill @ 10,589 tokens1,558.1 tok/s~1,725 tok/s
Prefill @ 42,550 tokens1,488.7 tok/s—
Decode, single stream (C1)69.59 tok/s~75.6 tok/s

So roughly 8-10% of throughput buys 147k additional KV tokens.

The prefill figures are conservative. The reference rig is power-capped at 300 W per GPU (1200 W total against a 600 W per-card maximum) and thermally throttles under sustained long-context prefill. An uncapped, better-cooled 4x RTX PRO 6000 box should measure higher prefill than published here. Decode is latency-bound on PCIe collectives and is much less sensitive to the cap.

Two traps worth stating plainly

`GRAPH` has a hard floor of `MAX_NUM_SEQS x (MTP + 1)` — 16 at SEQS=4, MTP=3. Setting GRAPH=8 buys 17,152 more KV tokens and collapses decode from 69.59 to 5.61 tok/s, because every decode step falls out of the captured full CUDA graph. Never set it below the decode row count.

`MAX_BATCHED_TOKENS` moves KV inversely. Transient buffers are max_batched-shaped and MTP-3 allocates two workspace lanes, doubling them. Raising 1,280 -> 3,072 cost 84,736 tokens. Raise it only when you want prefill throughput and are willing to pay KV for it.

Image requirement

This profile requires an image carrying the R7 fused MoE patches, e.g. verdictai/glm52-exl3-sparkinfer:v39-r28-r7fused-broadcast-cu132-sm120a (sha256:12f86065d7fe64d30dad678585e68c91f47f1f2a32bed45ccaf108382f3928ac). The public voipmonitor/vllm:infernal-invocation-...-20260817-r17 base ships exl3_moe_fused rather than the specialized exl3_moe_r7_fused symbol. The portable companion package adds the dedicated projection-mixed R7 ABI while preserving the legacy FP16 route ABI for ordinary EXL3 calls. That reconstructed image has not yet repeated the end-to-end GPU battery, so it is not the default. An older launcher routed around the available generic symbol and raised a missing-symbol error; the blanket claim that r17 cannot serve the checkpoint is superseded.

Repository map

text
calibration/
  reap_recall_calib.jsonl        exact live corpus
reproducibility/r10/
  README.md                       superseded B300 run and provenance
  RUN_METADATA.json               human-readable run parameters
  capture/                        exact prompt plan and capture summary
  inventories/                    source, carrier, numeric, runtime seals
  DEPLOYED_CODE.sha256             exact post-repair deployed-code hashes
  lineage/encode_tr3_v31.py       proven numeric core
  r7_encoder/                     exact deployed encoder package
  run/                            exact launch, supervisor, guard, assembly,
                                  finisher, and progressive-upload scripts
reproducibility/local-corrected-v1/
  README.md                       authoritative corrected reproduction guide
  STORAGE_BOUNDED_RUNBOOK.md      NVMe paging, seals, and reclamation order
  code/                           exact executed local correction code
  decisions/                      frozen decisions for all 75 routed layers
  receipts/                       preflight, source windows, and 75 seals
  tests/                          local correction regression tests
  results/                        corrected five-run KLD record
r7-experts-layer-*.safetensors    progressively uploaded expert weights
r7-experts-layer-*.json           per-layer bit maps and provenance

Loader compatibility boundary

This schema-v2 checkpoint requires a mixed-R7-aware loader. The pinned r34 and validated r17-derived runtimes described above implement that contract; stock or generic loaders without mixed K3/K4/K5 R7 support do not. The portable r17 companion implements the required dedicated ABI but remains a release candidate until its end-to-end replay is complete. The included converter refuses to label an unsupported conversion as ready.

Credits and upstream work

  • —Z.ai for GLM-5.2, the base model and license.
  • —Brandon Music for the owner corpus, TR3v4/R7 encoding and correction work, checkpoint construction, runtime integration, validation, and publication.
  • —turboderp for ExLlamaV3 and the EXL3 format and numeric core used here. EXL3 documents its lineage as a QTIP variant; see the EXL3 format note, QTIP, and QuIP#.
  • —malaiwah for the MTP-78 overlay and MTP-78 calibration capture.
  • —Josh Cartu for the MTP-78 recipe and rank-sliced runtime work credited by the upstream MTP overlay.
  • —Luke Alonso for the B12X kernels and Blackwell serving work on which these profiles depend.
  • —Martin Vit and yatesdr for the Infernal Invocation / RC2 image engineering credited by the upstream MTP overlay.
  • —Festr for isolating the shared-expert split-payload decode loss. This is a Discord handle; no public profile is linked because one has not been verified.
  • —Special thanks to the local-inference-lab community for its vLLM fork, Blackwell image work, review, testing, and release engineering.

The canonical contributor record is the GLM-5.2 TR3 quantization provenance.

The quantized weights inherit the license and use restrictions of the upstream GLM-5.2 model. ExLlamaV3 is MIT-licensed; consult every linked upstream project for its own license and terms. The proposed runtime bundle carries `THIRD_PARTY_NOTICES.md` and pinned Apache-2.0/MIT license copies alongside the redistributed patches.