mattbucci/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-AWQ
068
v3: populate config.json modules_to_not_convert + ignore with verified 14-entry list (was empty, caused SGLang to AWQ-wrap BF16 Mamba in_proj); README accuracy pass — actual 8-slice omni_thinking_tools recipe, drop_images=True approximation, verified serving on v0.5.13.post1+patches 052/053, 6/6 caps PASS, 256K decode 97.8 tok/s benchmarked. Weights byte-identical to v2.
v2: strip 5934 stale .weight_zero_point CT residue tensors (broke SGLang MoE loader); convert_moe_ct_to_awq.py patched upstream so future ships are clean by construction
AWQ 4-bit (group_size=64) of Nemotron-3-Nano-Omni-30B-A3B-Reasoning — calibrated 9d6h via llmcompressor GPTQ W4A16, repacked to native AWQ, scale audit clean against BF16 base
initial commit
