CoolFace
Modelpublic

Melvin56/DeepSeek-R1-Distill-Llama-8B-Enkrypt-Aligned-GGUF

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes103downloads
Model Card

Melvin56/DeepSeek-R1-Distill-Llama-8B-Enkrypt-Aligned-GGUF

Original Model : enkryptai/DeepSeek-R1-Distill-Llama-8B-Enkrypt-Aligned

All quants are made using the imatrix option.

ModelSize (GB)
Q2_K3.17
Q3KM4.02
Q4KM4.92
Q5KM5.72
Q6_K6.59
Q8_08.54
F1616.2
CPU (AVX2)CPU (ARM NEON)MetalcuBLASrocBLASSYCLCLBlastVulkanKompute
K-quants✅✅✅✅✅✅✅ 🐢5✅ 🐢5❌
I-quants✅ 🐢4✅ 🐢4✅ 🐢4✅✅Partial¹❌❌❌
✅: feature works
🚫: feature does not work
❓: unknown, please contribute if you can test it youself
🐢: feature is slow
¹: IQ3_S and IQ1_S, see #5886
²: Only with -ngl 0
³: Inference is 50% slower
⁴: Slower than K-quants of comparable size
⁵: Slower than cuBLAS/rocBLAS on similar cards
⁶: Only q8_0 and iq4_nl