Null-Guard/LFM2.5-230M-distilled-Gemini-3.8-Flash-Uncensored-GGUF
LFM2.5-230M-distilled-Gemini-3.8-Flash-Uncensored — GGUF
GGUF quantizations of Null-Guard/LFM2.5-230M-distilled-Gemini-3.8-Flash-Uncensored, a distilled, abliterated (refusal-suppressed) 230M-parameter model, for use with llama.cpp, Ollama, LM Studio, koboldcpp, and other GGUF-compatible runtimes.
⚠️ This model has had its safety alignment deliberately reduced via abliteration. Read Intended Use & Risks before using it.
Files
(Update this table with the actual quant files present in the repo)
Given the base model is only 230M parameters, even Q80 or F16 is small enough to run comfortably on CPU (and even on low-power/edge devices) — Q4KM or Q5KM is recommended for most users, with Q80/F16 as an option if you have the RAM/VRAM to spare and want maximum quality.
Quantization details
- Converted with:
llama.cpp(convert_hf_to_gguf.py+llama-quantize) — (fill in the exact commit/version you used) - Source weights: Null-Guard/LFM2.5-230M-distilled-Gemini-3.8-Flash-Uncensored (F32/F16 safetensors)
- Distillation: (fill in — describe the teacher/student distillation setup, data, and objective used to produce the base checkpoint)
- Imatrix used: (yes/no — if yes, note what calibration dataset was used for the importance matrix)
How to Use
llama.cpp
# Download a quant, e.g. Q4_K_M
huggingface-cli download Null-Guard/LFM2.5-230M-distilled-Gemini-3.8-Flash-Uncensored-GGUF \
lfm2.5-230m-uncensored.Q4_K_M.gguf --local-dir .
# Run with llama-cli
./llama-cli -m lfm2.5-230m-uncensored.Q4_K_M.gguf \
-p "You are a helpful assistant." \
-cnvOr serve it as an OpenAI-compatible API:
./llama-server -m lfm2.5-230m-uncensored.Q4_K_M.gguf -c 4096 --port 8080Ollama
Create a Modelfile:
FROM ./lfm2.5-230m-uncensored.Q4_K_M.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
"""
PARAMETER stop "<|im_end|>"Then:
ollama create lfm2.5-230m-uncensored -f Modelfile
ollama run lfm2.5-230m-uncensoredDouble-check the chat template above against the base model's actual template before publishing — copy it from the base model's tokenizer_config.json if it differs.LM Studio / koboldcpp / text-generation-webui
Any GGUF-compatible loader can load these files directly — search for Null-Guard/LFM2.5-230M-distilled-Gemini-3.8-Flash-Uncensored-GGUF in-app or point the loader at a downloaded .gguf file.
Choosing a Quant
- Q4_K_M — best default for most people; good balance of speed, size, and quality.
- Q5_K_M / Q6_K — if you want noticeably better output fidelity and can spare a bit more RAM.
- Q8_0 / F16 — if you want output as close as possible to the unquantized model (model is small enough that this is cheap).
- Q2_K / Q3_K_M — only if you are extremely constrained on RAM/storage; expect a real drop in coherence at this size, which will be more noticeable than on larger base models given the model only has 230M parameters to begin with.
Intended Use & Risks
This is a very small, permissive, uncensored, distilled model. Refusal behavior has been suppressed via abliteration on the base model before quantization — quantizing does not add or remove any safety behavior on its own.
- Intended for research, local experimentation, and personal/offline use.
- Not intended for public-facing deployment without your own moderation layer.
- Not intended for generating illegal content, sexual content involving minors, harassment, or other content prohibited by law in your jurisdiction — abliteration removes the model's tendency to refuse, it does not remove your responsibility for how you use the output.
- Not intended for use by minors.
- Because this checkpoint is distilled, its outputs may reflect the behavior, style, and any errors or biases of the teacher model it was distilled from — evaluate independently before relying on it for any downstream task.
Disclaimer: This model's safety filtering has been substantially reduced. It may produce inaccurate, biased, offensive, or otherwise harmful output. Use at your own risk; the maintainers of this repository provide it "as is" for research and personal use and do not endorse any specific downstream use.
Related
- Full-precision model: Null-Guard/LFM2.5-230M-distilled-Gemini-3.8-Flash-Uncensored
License
Apache 2.0. (Confirm this is compatible with the license terms of the distillation teacher model and any base architecture license before publishing — distilled models sometimes carry additional restrictions from the teacher model's terms of use.)
