brandonmusic/GLM-5.2-EXL3-TR3v4-3.5bpw-MTP78
docs: add reviewed attribution and credits
Correct r17 CUDA graph claim and mark eager launcher provisional
Make validated eager G64 runtime the download default
Fix r17 max-batched-token environment name
Make sealed r17 exact-G64-Q-only reader the download default
Document exact G64 Q-only five-run KLD pass
Document sealed r17 FP8 reader KLD comparison
Accuracy battery (LAVD/Estonia/Hotel/needle) on v39_502k, KV 502016
Add max-KV serving profile: 502,016 KV tokens @262K, MTP-3/DCP4 + prefill/decode results
Add max-KV serving profile: 502,016 KV tokens @262K, MTP-3/DCP4 + prefill/decode results
Add max-KV serving profile: 502,016 KV tokens @262K, MTP-3/DCP4 + prefill/decode results
Add max-KV serving profile: 502,016 KV tokens @262K, MTP-3/DCP4 + prefill/decode results
Model card: Gilded Gnosis r34 qualification — canonical run profile, release-gate results (credit: Festr and the local-inference-lab Discord)
Document and validate BF16-shared online K6 revision (#3)
Shared experts: BF16 + merged online K6 (split-payload decode loss isolated by Festr)
Model card: fused image published as verdictai/glm52-exl3-sparkinfer:v39-r28-r7fused-broadcast-cu132-sm120a
Point SERVING_FUSED.md at the published fused image
Point docker-compose.yaml at the published fused image
Point serve.sh at the published fused image
Model card: add fused MoE path (KV 355,328 @262K ctx, KLD 0.069527), document the required nvfp4 outer-scales file, and record the FC2-bound correctness fix
Add fused-MoE patches for custom image builds
Add fused-MoE patches for custom image builds
Add fused-MoE patches for custom image builds
Add fused-MoE patches for custom image builds
Add fused-MoE patches for custom image builds
Add fused-MoE patches for custom image builds
Add fused-MoE patches for custom image builds
Add SERVING_FUSED.md: fused MoE serving (KV 355,328 @262K ctx, KLD 0.069527)
Add docker-compose.yaml: fused MoE serving (KV 355,328 @262K ctx, KLD 0.069527)
Add serve.sh: fused MoE serving (KV 355,328 @262K ctx, KLD 0.069527)
Add nvfp4_mla_outer_scales.json: fused MoE serving (KV 355,328 @262K ctx, KLD 0.069527)
Model card: add measured prefill route-block-size tuning (+44.7% prefill, +21.8% decode) and the levers that did not help
SERVING.md: quote the measured KV capacity with its config instead of the auto-profile ceiling
Model card: KV cache capacities are now measured, with util/dtype/ctx/batch stated; qualify the 1.13M auto-profile figure
Model card: repair corrupted Run it block; link image, compose, serve script and results
Model card: surface Docker Hub image, compose, serve script and results up top
Model card: add KV cache section (supported dtypes + measured capacity) and link SERVING/RESULTS
Add RESULTS.md: serving (Docker Hub image + compose) and measured KLD/throughput results
Add SERVING.md: serving (Docker Hub image + compose) and measured KLD/throughput results
corrected r7 manifest: r7-experts-layer-077.json
corrected r7 experts: r7-experts-layer-077.safetensors
corrected r7 manifest: r7-experts-layer-076.json
corrected r7 experts: r7-experts-layer-076.safetensors
corrected r7 manifest: r7-experts-layer-075.json
corrected r7 experts: r7-experts-layer-075.safetensors
corrected r7 manifest: r7-experts-layer-074.json
corrected r7 experts: r7-experts-layer-074.safetensors
corrected r7 manifest: r7-experts-layer-073.json
corrected r7 experts: r7-experts-layer-073.safetensors
corrected r7 manifest: r7-experts-layer-072.json
