CoolFace
Modelpublic

neopolita/toolace-8b-gguf

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes95downloads
Model Card

GGUF quants for **Team-ACE/ToolACE-8B** using llama.cpp

Terms of Use: Please check the **original model**

<picture> <img alt="cthulhu" src="https://huggingface.co/neopolita/common/resolve/main/profile.png"> </picture>

Quants

  • —q2_k: Uses Q4K for the attention.vw and feedforward.w2 tensors, Q2_K for the other tensors.
  • —q3_k_s: Uses Q3_K for all tensors
  • —q3_k_m: Uses Q4K for the attention.wv, attention.wo, and feedforward.w2 tensors, else Q3_K
  • —q3_k_l: Uses Q5K for the attention.wv, attention.wo, and feedforward.w2 tensors, else Q3_K
  • —q4_0: Original quant method, 4-bit.
  • —q4_1: Higher accuracy than q40 but not as high as q50. However has quicker inference than q5 models.
  • —q4_k_s: Uses Q4_K for all tensors
  • —q4_k_m: Uses Q6K for half of the attention.wv and feedforward.w2 tensors, else Q4_K
  • —q5_0: Higher accuracy, higher resource usage and slower inference.
  • —q5_1: Even higher accuracy, resource usage and slower inference.
  • —q5_k_s: Uses Q5_K for all tensors
  • —q5_k_m: Uses Q6K for half of the attention.wv and feedforward.w2 tensors, else Q5_K
  • —q6_k: Uses Q8_K for all tensors
  • —q8_0: Almost indistinguishable from float16. High resource use and slow. Not recommended for most users.