CoolFace
Modelpublic

manalejandro/DeepSeek-R1-Distill-Qwen-1-5B-f16-iq3_xxs-GGUF

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes38downloads
Model Card

DeepSeek-R1-Distill-Qwen-1.5B (IQ3_XXS GGUF)

A heavily compressed GGUF quantization of DeepSeek-R1-Distill-Qwen-1.5B — the smallest official DeepSeek reasoning model.

This card covers the single file:

  • —`DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf`
PropertyValue
FileDeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf
Size733.44 MiB (769.07 MB)
QuantizationIQ3_XXS (≈ 3.06 bits/weight)
Parameters1.78B (total) / ~1.5B (active)
ArchitectureQwen2
Context length32768 tokens
GGUF version3 (latest)

Why this quantization

IQ3_XXS is the smallest IQ type that reliably completes the model's [Start thinking] / [End thinking] reasoning chain. At ~0.77 GB the weights load on any 4 GB-class GPU (e.g. GTX 1650) while still producing final answers instead of degenerating into an infinite thinking loop.

Note: even lighter quantizations exist (iq2_xxs at 2.06 bpw, ~0.56 GB) but on this reasoning model they frequently get stuck in the thinking loop and never emit the final answer.

It was produced from the F16 source with llama.cpp using an importance matrix computed with llama-imatrix over a wikitext-2 calibration set. The advertised context was capped at 32768 tokens so the KV cache fits a 4 GB GPU when fully offloaded (-ngl 999).

Usage

llama.cpp (CLI)

bash
llama-cli \
  -m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
  -p "What is the capital of France?" \
  -n 512 \
  -ngl 999 \
  -c 32768

llama-server (OpenAI-compatible API)

bash
llama-server \
  -m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
  -ngl 999 \
  -c 32768 \
  --port 8080
# curl http://localhost:8080/v1/chat/completions ...

Docker Model Runner (docker model)

bash
docker model package \
  --gguf DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf \
  --context-size 32768 \
  deepseek-r1
docker model run deepseek-r1 "What is the capital of France?"

Ollama

bash
ollama create deepseek-r1-iq3-xxs -f Modelfile   # FROM ./DeepSeek-R1-Distill-Qwen-1.5B-f16-iq3_xxs.gguf
ollama run deepseek-r1-iq3-xxs

Quality expectations

IQ3XXS is an aggressive compression, but it stays coherent enough to finish its reasoning and answer. Expect occasional arithmetic or factual slips and less fluency than F16 / Q4KM. Give the model enough output tokens: the DeepSeek-R1 chain-of-thought can use several hundred tokens before the answer (e.g. request `maxtokens >= 1024 from an API). For higher quality on the same hardware, prefer Q4KM (~1.0 GB); for a smaller file, IQ2_M` (~0.67 GB) also works but is slightly less reliable.

Credits & license