CoolFace
Modelpublic

manalejandro/DeepSeek-R1-Distill-Qwen-1-5B-f16-iq2_xxs-GGUF

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes43downloads
Model Card

DeepSeek-R1-Distill-Qwen-1.5B (IQ2_XXS GGUF)

A heavily compressed GGUF quantization of DeepSeek-R1-Distill-Qwen-1.5B — the smallest official DeepSeek reasoning model.

This card covers the single file:

  • —`DeepSeek-R1-Distill-Qwen-1.5B-f16-iq2_xxs.gguf`
PropertyValue
FileDeepSeek-R1-Distill-Qwen-1.5B-f16-iq2_xxs.gguf
Size560.37 MiB (587.59 MB)
QuantizationIQ2_XXS (≈ 2.06 bits/weight)
Parameters1.78B (total) / ~1.5B (active)
ArchitectureQwen2
Context length32768 tokens
GGUF version3 (latest)

Why this quantization

IQ2_XXS is the smallest IQ type that keeps the model usable on low-VRAM hardware. At ~0.56 GB the weights load on any 4 GB-class GPU (e.g. GTX 1650) while retaining the model's reasoning behaviour, including the [Start thinking] / [End thinking] chain-of-thought format.

It was produced from the F16 source with llama.cpp using an importance matrix computed with llama-imatrix over a wikitext-2 calibration set. The advertised context was capped at 32768 tokens so the KV cache fits a 4 GB GPU when fully offloaded (-ngl 999).

Usage

llama.cpp (CLI)

bash
llama-cli \
  -m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq2_xxs.gguf \
  -p "What is the capital of France?" \
  -n 256 \
  -ngl 999 \
  -c 32768

llama-server (OpenAI-compatible API)

bash
llama-server \
  -m DeepSeek-R1-Distill-Qwen-1.5B-f16-iq2_xxs.gguf \
  -ngl 999 \
  -c 32768 \
  --port 8080
# curl http://localhost:8080/v1/chat/completions ...

Docker Model Runner (docker model)

bash
docker model package \
  --gguf DeepSeek-R1-Distill-Qwen-1.5B-f16-iq2_xxs.gguf \
  --context-size 32768 \
  deepseek-r1
docker model run deepseek-r1 "What is the capital of France?"

Ollama

bash
ollama create deepseek-r1-iq2-xxs -f Modelfile   # FROM ./DeepSeek-R1-Distill-Qwen-1.5B-f16-iq2_xxs.gguf
ollama run deepseek-r1-iq2-xxs

Quality expectations

IQ2XXS is an aggressive compression. The model answers factual questions correctly and keeps its reasoning structure, but expect degraded fluency and precision compared to F16 / Q4KM. For higher quality on the same hardware, prefer `Q4KM` (~1.0 GB); for absolute minimum size, `IQ1S or Q1_0` exist but degrade much faster.

Credits & license