CoolFace
Modelpublic

mattbucci/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-AWQ

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes68downloads
4 commits on main
87b1dda3mo ago

v3: populate config.json modules_to_not_convert + ignore with verified 14-entry list (was empty, caused SGLang to AWQ-wrap BF16 Mamba in_proj); README accuracy pass — actual 8-slice omni_thinking_tools recipe, drop_images=True approximation, verified serving on v0.5.13.post1+patches 052/053, 6/6 caps PASS, 256K decode 97.8 tok/s benchmarked. Weights byte-identical to v2.

mattbucci
4c987113mo ago

v2: strip 5934 stale .weight_zero_point CT residue tensors (broke SGLang MoE loader); convert_moe_ct_to_awq.py patched upstream so future ships are clean by construction

mattbucci
9edacec3mo ago

AWQ 4-bit (group_size=64) of Nemotron-3-Nano-Omni-30B-A3B-Reasoning — calibrated 9d6h via llmcompressor GPTQ W4A16, repacked to native AWQ, scale audit clean against BF16 base

mattbucci
bc9eb793mo ago

initial commit

mattbucci