malaiwah/GLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing) Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048 tokens of the exllamav3 standardcaldata corpus (pinned): per context, layer_NNN.attn_in and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlpin is the router + expert gate/up input) and `layerNNN.router_logits` (fp32, natural top-8 routing ground truth). Per-expert Hessians E[xx^T], routing statistics and down-proj inputs are recomputable offline. Captured with vLLM TP8 eager; manifest carries sha256 per file.
Note: these activations are one draw from the engine-launch distribution — the runtime is not launch-deterministic (Triton autotune winner selection on the KDA kernels; see the fidelity suite's nondeterminism receipts and https://github.com/vllm-project/vllm/pull/53906#issuecomment-5433635837 ). For Hessian/routing statistics over 188K tokens this launch noise is far below calibration sampling noise. A pinned-env (TRITONCACHEAUTOTUNING) v2 recapture is scripted in the companion repo for bit-reproducible needs.
