manalejandro/DeepSeek-R1-Distill-Qwen-1-5B-f16-iq3_xxs-GGUF
DeepSeek-R1-Distill-Qwen-1.5B (IQ3_XXS GGUF)
A heavily compressed GGUF quantization of DeepSeek-R1-Distill-Qwen-1.5B — the smallest official DeepSeek reasoning model.
This card covers the single file:
- `DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf`
Why this quantization
IQ3_XXS is the smallest IQ type that reliably completes the model's [Start thinking] / [End thinking] reasoning chain. At ~0.77 GB the weights load on any 4 GB-class GPU (e.g. GTX 1650) while still producing final answers instead of degenerating into an infinite thinking loop.
Note: even lighter quantizations exist (iq2_xxs at 2.06 bpw, ~0.56 GB) but on this reasoning model they frequently get stuck in the thinking loop and never emit the final answer.It was produced from the F16 source with llama.cpp using an importance matrix computed with llama-imatrix over a wikitext-2 calibration set. The advertised context was capped at 32768 tokens so the KV cache fits a 4 GB GPU when fully offloaded (-ngl 999).
Usage
llama.cpp (CLI)
llama-cli \
-m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
-p "What is the capital of France?" \
-n 512 \
-ngl 999 \
-c 32768llama-server (OpenAI-compatible API)
llama-server \
-m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
-ngl 999 \
-c 32768 \
--port 8080
# curl http://localhost:8080/v1/chat/completions ...Docker Model Runner (docker model)
docker model package \
--gguf DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
--context-size 32768 \
deepseek-r1
docker model run deepseek-r1 "What is the capital of France?"Ollama
ollama create deepseek-r1-iq3-xxs -f Modelfile # FROM ./DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf
ollama run deepseek-r1-iq3-xxsQuality expectations
IQ3XXS is an aggressive compression, but it stays coherent enough to finish its reasoning and answer. Expect occasional arithmetic or factual slips and less fluency than F16 / Q4KM. Give the model enough output tokens: the DeepSeek-R1 chain-of-thought can use several hundred tokens before the answer (e.g. request `maxtokens >= 1024 from an API). For higher quality on the same hardware, prefer Q4KM (~1.0 GB); for a smaller file, IQ2_M` (~0.67 GB) also works but is slightly less reliable.
Credits & license
- Base model: deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B (MIT)
- Source GGUF: bartowski/DeepSeek-R1-Distill-Qwen-1.5B-GGUF (MIT)
- Quantized with llama.cpp +
llama-imatrix - This file is released under the MIT license.
