qwp4w3hyb/Hunyuan-A13B-Instruct-hf-WIP-GGUF
2358
Alpha quality, needs WIP PR [outdated]
- The quants here are outdated, bullerwins/Hunyuan-A13B-Instruct-GGUF includes more recent fixes and will likely work better.
- Based on ngxson/llama.cpp/pull/26@46c8b70cbc7346db95e45ebae4f1e0c68a9b8d86
- which is based on ggml-org/llama.cpp/pull/14425
- supposedly works mostly fine™ when run with below args according to ggml-org/llama.cpp/pull/14425#issuecomment-3017533726
--ctx-size 262144 -b 1024 --jinja --no-warmup --cache-type-k q8_0 --cache-type-v q8_0 --flash-attn --temp 0.6 --presence-penalty 0.7 --min-p 0.1- might still be very broken, no guarantees!
