MikeKuykendall/deepseek-moe-16b-q8-0-cpu-offload-gguf
0107
DeepSeek-MoE-16B Q8_0 with CPU Offloading
Q8_0 quantization of DeepSeek-MoE-16B with CPU offloading support. Highest quality, near-F16 accuracy.
Performance
File Size: 17 GB (from 31 GB F16)
Usage
huggingface-cli download MikeKuykendall/deepseek-moe-16b-q8-0-cpu-offload-gguf
shimmy serve --model-dirs ./models --cpu-moeLinks: Q2_K | Q4_K_M
License: Apache 2.0
