CoolFace
Modelpublic

minjaechoi/nemotron3-nano-30b-a3b-gsq-2p25bit-r35

sourceHugging Faceupdated 6d agoView on Hugging Face
0likes161downloads
Model Card

NVIDIA-Nemotron-3-Nano-30B-A3B -- 2-bit routed experts (r35)

Internal research checkpoint. Routed experts average 2.25 bits; every other weight is BF16. Weights are stored dequantized in BF16 tensors and load with stock transformers / vLLM.

GSQ baseline (group size 64, 2-bit packed codes with bf16 group scales), dequantized to BF16. This is the GSQ row of the Nemotron block in the main table.

Routed-expert average2.25 bits
Internal IDr35

License follows the base model.