CoolFace
Modelpublic

CCATresearch/Gemma-2-2B_wllama_gguf

sourceHugging Facegemmaupdated 2y agoView on Hugging Face
0likes55downloads
Model Card

Gemma 2 2B quantized for wllama (under 2gb).

q4048 is WAY faster when using llama.cpp, with wllama, it's about the same as q4k.