CoolFace
Modelpublic

about0/llama-v2-chinese-ymcui-alpaca-GGML-13B

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes
Model Card

LLaMA-v2-chinese-alpaca-13B-GGML (ymcui)

Here are the GGML converted and/or quantized models for ymcui's Chinese LLaMA-v2 Alpaca 13B.

!NOTE! The GGML filetype is outdated. Prefer GGUF format going forward.

Explanation of quantisation methods

<details> <summary>Click to see details</summary>

Methods:

  • —type-0 (Q40, Q50, Q8_0) - weights w are obtained from quants q using w = d * q, where d is the block scale.
  • —type-1 (Q41, Q51) - weights are given by w = d * q + m, where m is the block minimum

The new methods available are:

  • —GGMLTYPEQ2_K - "type-1" 2-bit quantization in super-blocks containing 16 blocks, each block having 16 weight. Block scales and mins are quantized with 4 bits. This ends up effectively using 2.5625 bits per weight (bpw)
  • —GGMLTYPEQ3_K - "type-0" 3-bit quantization in super-blocks containing 16 blocks, each block having 16 weights. Scales are quantized with 6 bits. This end up using 3.4375 bpw.
  • —GGMLTYPEQ4_K - "type-1" 4-bit quantization in super-blocks containing 8 blocks, each block having 32 weights. Scales and mins are quantized with 6 bits. This ends up using 4.5 bpw.
  • —GGMLTYPEQ5K - "type-1" 5-bit quantization. Same super-block structure as GGMLTYPEQ4K resulting in 5.5 bpw
  • —GGMLTYPEQ6_K - "type-0" 6-bit quantization. Super-blocks with 16 blocks, each block having 16 weights. Scales are quantized with 8 bits. This ends up using 6.5625 bpw
  • —GGMLTYPEQ8K - "type-0" 8-bit quantization. Only used for quantizing intermediate results. The difference to the existing Q80 is that the block size is 256. All 2-6 bit dot products are implemented for this quantization type.

This is exposed via llama.cpp quantization types that define various "quantization mixes" as follows:

  • —LLAMAFTYPEMOSTLYQ2K - uses GGMLTYPEQ4K for the attention.vw and feedforward.w2 tensors, GGMLTYPEQ2_K for the other tensors.
  • —LLAMAFTYPEMOSTLYQ3KS - uses GGMLTYPEQ3K for all tensors
  • —LLAMAFTYPEMOSTLYQ3KM - uses GGMLTYPEQ4K for the attention.wv, attention.wo, and feedforward.w2 tensors, else GGMLTYPEQ3K
  • —LLAMAFTYPEMOSTLYQ3KL - uses GGMLTYPEQ5K for the attention.wv, attention.wo, and feedforward.w2 tensors, else GGMLTYPEQ3K
  • —LLAMAFTYPEMOSTLYQ4KS - uses GGMLTYPEQ4K for all tensors
  • —LLAMAFTYPEMOSTLYQ4KM - uses GGMLTYPEQ6K for half of the attention.wv and feedforward.w2 tensors, else GGMLTYPEQ4K
  • —LLAMAFTYPEMOSTLYQ5KS - uses GGMLTYPEQ5K for all tensors
  • —LLAMAFTYPEMOSTLYQ5KM - uses GGMLTYPEQ6K for half of the attention.wv and feedforward.w2 tensors, else GGMLTYPEQ5K
  • —LLAMAFTYPEMOSTLYQ6K- uses 6-bit quantization (GGMLTYPEQ8_K) for all tensors </details>

Provided files

NameQuant methodBitsSizeMax RAM requiredUse case
llama-v2-chinese-alpaca-13B-Q2_K.ggmlQ2_K25.65 GB8.15 GBsmallest, significant quality-loss - not recommended for most purposes
llama-v2-chinese-alpaca-13B-Q3_K_S.ggmlQ3KS35.81 GB8.31 GBvery small, high quality-loss
llama-v2-chinese-alpaca-13B-Q3_K_M.ggmlQ3KM36.46 GB7.96 GBvery small, high quality-loss
llama-v2-chinese-alpaca-13B-Q3_K_L.ggmlQ3KL37.08 GB9.58 GBsmall, substantial quality-loss
llama-v2-chinese-alpaca-13B-Q4_0.ggmlQ4_047.53 GB10.03 GBlegacy; small, very high quality-loss - prefer using Q3KM
llama-v2-chinese-alpaca-13B-Q4_1.ggmlQ4_148.34 GB10.84 GBlegacy; small, very high quality-loss - prefer using Q3KM
llama-v2-chinese-alpaca-13B-Q4_K_S.ggmlQ4KS47.53 GB10.03 GBsmall, greater quality-loss
llama-v2-chinese-alpaca-13B-Q4_K_M.ggmlQ4KM48.03 GB10.53 GBmedium, balanced quality - recommended
llama-v2-chinese-alpaca-13B-Q5_0.ggmlQ5_059.15 GB11.65 GBlegacy; medium, balanced quality - prefer using Q4KM
llama-v2-chinese-alpaca-13B-Q5_1.ggmlQ5_159.96 GB12.46 GBlegacy; medium, balanced quality - prefer using Q4KM
llama-v2-chinese-alpaca-13B-Q5_K_S.ggmlQ5KS59.15 GB11.65 GBlarge, low quality-loss - recommended
llama-v2-chinese-alpaca-13B-Q5_K_M.ggmlQ5KM59.41 GB11.91 GBlarge, very low quality-loss - recommended
llama-v2-chinese-alpaca-13B-Q6_K.ggmlQ6_K610.9 GB13.4 GBvery large, extremely low quality-loss
llama-v2-chinese-alpaca-13B-Q8_0.ggmlQ8_0814 GB16.5 GBvery large, extremely low quality-loss - not recommended
llama-v2-chinese-alpaca-13B-f16.ggmlf161626.5 GB29 GBvery large, almost no quality-loss - not recommended

Model Sources

  • —Repository: [https://github.com/ymcui/Chinese-LLaMA-Alpaca-2]