inference-snaps/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-5GB
1310
Nemotron 3.5 Lightning 30B A3B 5GB
Q4KM
The NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 model, quantized to Q4KM and split into 5GB GGUF files.
The model has been split using llama-gguf-split (b10237) as follows:
llama-gguf-split --split-max-size 5G NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_MNVFP4
The NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 model, split into 5GB GGUF files.
llama-gguf-split --split-max-size 5G NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4.gguf NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4