CoolFace
Modelpublic

inference-snaps/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-5GB

sourceHugging Faceupdated 6d agoView on Hugging Face
1likes310downloads
Model Card

Nemotron 3.5 Lightning 30B A3B 5GB

Q4KM

The NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 model, quantized to Q4KM and split into 5GB GGUF files.

The model has been split using llama-gguf-split (b10237) as follows:

shell
llama-gguf-split --split-max-size 5G NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M

NVFP4

The NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model, split into 5GB GGUF files.

shell
llama-gguf-split --split-max-size 5G NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4.gguf NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4