chrispie/llama-hqq-1-bit
0
Demo for HQQ 1-bit quantized (binary weights) Llama2-7B-chat model using a low-rank adapter to improve the performance (referred to as HQQ+). You will need a GPU for this.
https://huggingface.co/mobiuslabsgmbh/Llama-2-7b-chat-hf1bitgs8hqq
