CoolFace
Modelpublic

Null-Guard/Qwen3-0.6B-Uncensored-GGUF

sourceHugging Faceapache-2.0updated 27d agoView on Hugging Face
1likes1.5kdownloads
Model Card

Qwen3-0.6B-Uncensored — GGUF

GGUF quantizations of Null-Guard/Qwen3-0.6B-Uncensored, an abliterated (refusal-suppressed) version of Qwen/Qwen3-0.6B, for use with llama.cpp, Ollama, LM Studio, koboldcpp, and other GGUF-compatible runtimes.

⚠️ This model has had its safety alignment deliberately reduced via abliteration. Read Intended Use & Risks before using it.

Files

(Update this table with the actual quant files present in the repo)

FilenameQuant typeSizeNotes
qwen3-0.6b-uncensored.Q2_K.ggufQ2_K~0.3 GBSmallest, noticeable quality loss
qwen3-0.6b-uncensored.Q3_K_M.ggufQ3KM~0.35 GBLow resource use
qwen3-0.6b-uncensored.Q4_K_M.ggufQ4KM~0.4 GBRecommended balance of size/quality
qwen3-0.6b-uncensored.Q5_K_M.ggufQ5KM~0.45 GBBetter quality, still small
qwen3-0.6b-uncensored.Q6_K.ggufQ6_K~0.5 GBNear-lossless
qwen3-0.6b-uncensored.Q8_0.ggufQ8_0~0.65 GBHighest quality quant, largest size
qwen3-0.6b-uncensored.f16.ggufF16~1.2 GBFull precision, for re-quantizing

Given the base model is only 0.6B parameters, even Q80 or F16 is small enough to run comfortably on CPU — Q4KM or Q5KM is recommended for most users, with Q80/F16 as an option if you have the RAM/VRAM to spare and want maximum quality.

Quantization details

  • —Converted with: llama.cpp (convert_hf_to_gguf.py + llama-quantize) — (fill in the exact commit/version you used)
  • —Source weights: Null-Guard/Qwen3-0.6B-Uncensored (F32/F16 safetensors)
  • —Imatrix used: (yes/no — if yes, note what calibration dataset was used for the importance matrix)

How to Use

llama.cpp

bash
# Download a quant, e.g. Q4_K_M
huggingface-cli download Null-Guard/Qwen3-0.6B-Uncensored-GGUF \
  qwen3-0.6b-uncensored.Q4_K_M.gguf --local-dir .

# Run with llama-cli
./llama-cli -m qwen3-0.6b-uncensored.Q4_K_M.gguf \
  -p "You are a helpful assistant." \
  -cnv

Or serve it as an OpenAI-compatible API:

bash
./llama-server -m qwen3-0.6b-uncensored.Q4_K_M.gguf -c 4096 --port 8080

Ollama

Create a Modelfile:

FROM ./qwen3-0.6b-uncensored.Q4_K_M.gguf

TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
"""

PARAMETER stop "<|im_end|>"

Then:

bash
ollama create qwen3-0.6b-uncensored -f Modelfile
ollama run qwen3-0.6b-uncensored
Double-check the chat template above against Qwen3's actual template before publishing — copy it from the base model's tokenizer_config.json if it differs.

LM Studio / koboldcpp / text-generation-webui

Any GGUF-compatible loader can load these files directly — search for Null-Guard/Qwen3-0.6B-Uncensored-GGUF in-app or point the loader at a downloaded .gguf file.

Choosing a Quant

  • —Q4_K_M — best default for most people; good balance of speed, size, and quality.
  • —Q5_K_M / Q6_K — if you want noticeably better output fidelity and can spare a bit more RAM.
  • —Q8_0 / F16 — if you want output as close as possible to the unquantized model (model is small enough that this is cheap).
  • —Q2_K / Q3_K_M — only if you are extremely constrained on RAM/storage; expect a real drop in coherence at this size.

Intended Use & Risks

This is a small, permissive, uncensored model. Refusal behavior has been suppressed via abliteration on the base model before quantization — quantizing does not add or remove any safety behavior on its own.

  • —Intended for research, local experimentation, and personal/offline use.
  • —Not intended for public-facing deployment without your own moderation layer.
  • —Not intended for generating illegal content, sexual content involving minors, harassment, or other content prohibited by law in your jurisdiction — abliteration removes the model's tendency to refuse, it does not remove your responsibility for how you use the output.
  • —Not intended for use by minors.

Disclaimer: This model's safety filtering has been substantially reduced. It may produce inaccurate, biased, offensive, or otherwise harmful output. Use at your own risk; the maintainers of this repository provide it "as is" for research and personal use and do not endorse any specific downstream use.

Related

License

Apache 2.0, inherited from Qwen3-0.6B. See the base model's license for full terms.