CoolFace
Modelpublic

constructai/VibeThinker-1.5B-GGUF

sourceHugging Facemitupdated 3mo agoView on Hugging Face
0likes808downloads
Model Card

constructai/VibeThinker-1.5B-GGUF

This is a quantized version of the original VibeThinker-1.5B , converted to the GGUF format for efficient CPU/GPU inference with llama.cpp, Ollama, or any GGUF‑compatible runner.


Original Model


Available Quantizations

Choose the quantization that fits your needs:

QuantizationFile Size
UD-IQ1_S437 MB
UD-IQ1_M464 MB
UD-IQ2_XXS511 MB
Q2_K676 MB
UD-IQ2_M601 MB
UD-Q2_K_XL676 MB
UD-IQ3_XXS669 MB
Q3_K_S761 MB
UD-IQ3_S762 MB
Q3_K_M824 MB
UD-Q3_K_M824 MB
UD-Q3_K_XL880 MB
UD-IQ4_XS896 MB
Q4_K_S940 MB
UD-IQ4_NL936 MB
Q4_K_M986 MB
UD-Q4_K_XL986 MB
Q5_K_S1.1 GB
UD-Q5_K_S1.1 GB
Q5_K_M1.13 GB
UD-Q5_K_M1.13 GB
UD-Q5_K_XL1.13 GB
Q6_K1.27 GB
UD-Q6_K1.27 GB
UD-Q6_K_XL1.27 GB
Q8_01.65 GB
UD-Q8_K_XL1.65 GB
F163.09 GB

For a 1.5B‑parameter model, even the larger files are quite manageable. Here’s what I recommend: F16 (3.09 GB) or Q8_0 (1.65 GB).

The other quants are also usable!


Usage

With ollama

bash
ollama run hf.co/constructai/VibeThinker-1.5B-GGUF:F16

With llama.cpp

bash
llama-server -hf constructai/VibeThinker-1.5B-GGUF:VibeThinker-1.5B-GGUF-F16.gguf

or

bash
llama-cli -hf constructai/VibeThinker-1.5B-GGUF:VibeThinker-1.5B-GGUF-F16.gguf

With LM Studio

bash
lms get constructai/VibeThinker-1.5B-GGUF@F16