MikeKuykendall/deepseek-moe-16b-q2-k-cpu-offload-gguf
047
DeepSeek-MoE-16B Q2_K with CPU Offloading
Q2_K quantization of DeepSeek-MoE-16B with CPU offloading support. Smallest size, maximum VRAM savings.
Performance
File Size: 6.3 GB (from 31 GB F16)
Usage
huggingface-cli download MikeKuykendall/deepseek-moe-16b-q2-k-cpu-offload-gguf
shimmy serve --model-dirs ./models --cpu-moeLinks: Q4_K_M | Q8_0
License: Apache 2.0
