CoolFace
Modelpublic

MikeKuykendall/deepseek-moe-16b-q8-0-cpu-offload-gguf

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes107downloads
Model Card

DeepSeek-MoE-16B Q8_0 with CPU Offloading

Q8_0 quantization of DeepSeek-MoE-16B with CPU offloading support. Highest quality, near-F16 accuracy.

Performance

ConfigurationVRAMSavedReduction
All GPU17.11 GB--
CPU Offload2.33 GB14.78 GB86.4%

File Size: 17 GB (from 31 GB F16)

Usage

bash
huggingface-cli download MikeKuykendall/deepseek-moe-16b-q8-0-cpu-offload-gguf
shimmy serve --model-dirs ./models --cpu-moe

Links: Q2_K | Q4_K_M

License: Apache 2.0