jayzou3773/less-is-moe-qwen3.5-122b-a10b-s1-128-seq8192-intdim-l-50
IntDim-L 50% pruned Qwen/Qwen3.5-122B-A10B
This checkpoint was structurally pruned with the released Less-is-MoE mean-absolute-gradient method. It removes exactly 50% of routed-expert FFN neurons using 128 calibration samples from yentinglin/s1K-1.1-trl-format revision 58a01564d278477da20ead1bcf1cde8e31f36251. The released loader settings are preserved: train, messages, shuffle_seed=1234, seq_length=8192, prefix truncation, no padding, and BF16, with no optimizer step. The exact model-specific token tensors are published at jayzou3773/less-is-moe-s1-calibration-128-seq8192 revision 678b4e666183e16ec00376960df03b6381632ed1. The source checkpoint was loaded and pruned in BF16.
The source-row selection hash is f261e952d4e6d5dec6d37db4ab22636761b215281d33896060fe4467bb352784 and the model-specific token-file hash is 47214818e5c0acaa6e65d3f212f7fab62937a4085d26394c356e3be1c94fffbf. Full export and zero-mask equivalence metadata are in experiment-export.json.
Inference requires the Less-is-MoE ragged vLLM plugin from the unified Less-is-MoE GPU image. IntDim-E has one uniform expert width. IntDim-L/G retain the routed MoE topology and store compact per-expert widths in config.json.
