Null-Guard/Qwen3-0.6B-Uncensored-GGUF
Qwen3-0.6B-Uncensored — GGUF
GGUF quantizations of Null-Guard/Qwen3-0.6B-Uncensored, an abliterated (refusal-suppressed) version of Qwen/Qwen3-0.6B, for use with llama.cpp, Ollama, LM Studio, koboldcpp, and other GGUF-compatible runtimes.
⚠️ This model has had its safety alignment deliberately reduced via abliteration. Read Intended Use & Risks before using it.
Files
(Update this table with the actual quant files present in the repo)
Given the base model is only 0.6B parameters, even Q80 or F16 is small enough to run comfortably on CPU — Q4KM or Q5KM is recommended for most users, with Q80/F16 as an option if you have the RAM/VRAM to spare and want maximum quality.
Quantization details
- Converted with:
llama.cpp(convert_hf_to_gguf.py+llama-quantize) — (fill in the exact commit/version you used) - Source weights: Null-Guard/Qwen3-0.6B-Uncensored (F32/F16 safetensors)
- Imatrix used: (yes/no — if yes, note what calibration dataset was used for the importance matrix)
How to Use
llama.cpp
# Download a quant, e.g. Q4_K_M
huggingface-cli download Null-Guard/Qwen3-0.6B-Uncensored-GGUF \
qwen3-0.6b-uncensored.Q4_K_M.gguf --local-dir .
# Run with llama-cli
./llama-cli -m qwen3-0.6b-uncensored.Q4_K_M.gguf \
-p "You are a helpful assistant." \
-cnvOr serve it as an OpenAI-compatible API:
./llama-server -m qwen3-0.6b-uncensored.Q4_K_M.gguf -c 4096 --port 8080Ollama
Create a Modelfile:
FROM ./qwen3-0.6b-uncensored.Q4_K_M.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
"""
PARAMETER stop "<|im_end|>"Then:
ollama create qwen3-0.6b-uncensored -f Modelfile
ollama run qwen3-0.6b-uncensoredDouble-check the chat template above against Qwen3's actual template before publishing — copy it from the base model's tokenizer_config.json if it differs.LM Studio / koboldcpp / text-generation-webui
Any GGUF-compatible loader can load these files directly — search for Null-Guard/Qwen3-0.6B-Uncensored-GGUF in-app or point the loader at a downloaded .gguf file.
Choosing a Quant
- Q4_K_M — best default for most people; good balance of speed, size, and quality.
- Q5_K_M / Q6_K — if you want noticeably better output fidelity and can spare a bit more RAM.
- Q8_0 / F16 — if you want output as close as possible to the unquantized model (model is small enough that this is cheap).
- Q2_K / Q3_K_M — only if you are extremely constrained on RAM/storage; expect a real drop in coherence at this size.
Intended Use & Risks
This is a small, permissive, uncensored model. Refusal behavior has been suppressed via abliteration on the base model before quantization — quantizing does not add or remove any safety behavior on its own.
- Intended for research, local experimentation, and personal/offline use.
- Not intended for public-facing deployment without your own moderation layer.
- Not intended for generating illegal content, sexual content involving minors, harassment, or other content prohibited by law in your jurisdiction — abliteration removes the model's tendency to refuse, it does not remove your responsibility for how you use the output.
- Not intended for use by minors.
Disclaimer: This model's safety filtering has been substantially reduced. It may produce inaccurate, biased, offensive, or otherwise harmful output. Use at your own risk; the maintainers of this repository provide it "as is" for research and personal use and do not endorse any specific downstream use.
Related
- Full-precision model: Null-Guard/Qwen3-0.6B-Uncensored
- Base model: Qwen/Qwen3-0.6B
License
Apache 2.0, inherited from Qwen3-0.6B. See the base model's license for full terms.
