lalitkarthik/scholar-3b-gguf
010
scholar-3b — GGUF
Q4KM quantisation of lalitkarthik/scholar-3b-v2 merged into Qwen/Qwen2.5-3B-Instruct, for CPU inference with llama.cpp.
Judged win rate against the base model: 85.7% over 60 held-out prompts.
llama-server --hf-repo lalitkarthik/scholar-3b-gguf --hf-file scholar-3b-Q4_K_M.gguf