CoolFace
Apppublic

chrispie/llama-hqq-1-bit

sourceHugging Facellama2updated 2y agoView on Hugging Face
0likes
App README

Demo for HQQ 1-bit quantized (binary weights) Llama2-7B-chat model using a low-rank adapter to improve the performance (referred to as HQQ+). You will need a GPU for this.

https://huggingface.co/mobiuslabsgmbh/Llama-2-7b-chat-hf1bitgs8hqq