CoolFace
Modelpublic

mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q4_K_M-GGUF

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
1likes1.2kdownloads
Model Card

Nemotron-3.5-Lightning-30B-A3B Heretic-Abliterated (Q4KM GGUF)

GGUF quantization of mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16 — NVIDIA-Nemotron-3.5-Lightning-30B-A3B (31.6B total / 3B active) with its refusal direction removed via Heretic.

What this is for: the same abliterated hybrid Mamba-MoE model, quantized to 24.3 GB so it runs locally via llama.cpp.

  • —Architecture: nemotron_h_moe (requires llama.cpp build b10326+)
  • —Quantization: Q4KM
  • —File size: 24.3 GB
  • —Smoke-tested locally before upload (loads + coherent output on llama-cli).

Results

RefusalsComplianceKL DivergenceTrials
0%100%0.0397200

Independent eval of the merged BF16 model (50 harmful-behavior prompts). The automated Zou keyword detector false-positives on words like "illegal"/"unethical" appearing inside compliant answers; manual review found 0 genuine refusals.

Usage

bash
llama-cli -m Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q4_K_M-GGUF.gguf \
  -cnv -st -p "What is 2+2?"

Ollama

Create a Modelfile:

dockerfile
FROM ./Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q4_K_M-GGUF.gguf

Then:

bash
ollama create nemotron-3.5-30b-heretic-q4_k_m
ollama run nemotron-3.5-30b-heretic-q4_k_m
Abliteration removes safety alignment. Use responsibly and in accordance with your local laws and the upstream NVIDIA Open Model License.