mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q6_K-GGUF
1346
Nemotron-3.5-Lightning-30B-A3B Heretic-Abliterated (Q6_K GGUF)
GGUF quantization of mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16 — NVIDIA-Nemotron-3.5-Lightning-30B-A3B (31.6B total / 3B active) with its refusal direction removed via Heretic.
What this is for: the same abliterated hybrid Mamba-MoE model, quantized to 33.5 GB so it runs locally via llama.cpp.
- Architecture:
nemotron_h_moe(requires llama.cpp build b10326+) - Quantization: Q6_K
- File size: 33.5 GB
- Smoke-tested locally before upload (loads + coherent output on
llama-cli).
Results
Independent eval of the merged BF16 model (50 harmful-behavior prompts). The automated Zou keyword detector false-positives on words like "illegal"/"unethical" appearing inside compliant answers; manual review found 0 genuine refusals.
Usage
llama-cli -m Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q6_K-GGUF.gguf \
-cnv -st -p "What is 2+2?"Ollama
Create a Modelfile:
FROM ./Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q6_K-GGUF.ggufThen:
ollama create nemotron-3.5-30b-heretic-q6_k
ollama run nemotron-3.5-30b-heretic-q6_kAbliteration removes safety alignment. Use responsibly and in accordance with your local laws and the upstream NVIDIA Open Model License.
