darioooooo0o/K2-Horizon-7B-GGUF
0750
K2-Horizon-7B GGUF quants

Requests, questions or suggestions? Message me on X: https://x.com/imdariotoo
GGUF quantizations of IFM/K2-Horizon-7B.
Converted with the official k2-official llama.cpp branch (MBZUAI-IFM port, commit 35999d101).
Files
Requirements
Use a llama.cpp build from the `k2-official` branch of MBZUAI-IFM/llama.cpp (or anything that merges that port). Mainline llama.cpp does NOT support the k2-horizon architecture.
Usage
llama-cli -m k2horizon7-q4_k_m.gguf -ngl 99 -c 8192Fits entirely on an 8 GB+ GPU; no CPU offload needed.
Notes
- Plain K-quants from BF16, no imatrix.
- All quants verified loading and generating on RTX 3060 12GB.
