CoolFace
Modelpublic

ChenMnZ/Llama-3-8b-BlockAP-w3g128

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes13downloads
Model Card

Block-AP (EfficientQAT w/o E2E-AP)

EfficientQAT involves two consecutive training phases: Block-wise training of all parameters (Block-AP) and end-to-end training of quantization parameters (E2E-QP).

In this repo, we provide the quantized checkpoints of Block-AP. Anyone can use them to reproduce our results or carry following research.

Performance

ModelQuantizationWikiText2 PPLAvg. AccuracyModel Size (GB)Hub link
Llama-2-7Bfp165.4764.8613.2-
Llama-2-7Bw4g1285.5664.073.7Link
Llama-2-7Bw3g1285.8963.963.1Link
Llama-2-7Bw2g647.6559.542.3Link
Llama-2-7Bw2g1287.9458.722.2Link
Llama-2-13Bfp164.8867.8125.4-
Llama-2-13Bw4g1284.9667.276.8Link
Llama-2-13Bw3g1285.2067.305.6Link
Llama-2-13Bw2g646.5563.104.0Link
Llama-2-13Bw2g1286.6863.493.8Link
Llama-2-70Bfp163.3272.41131.6-
Llama-2-70Bw4g1283.4172.5435.8Link
Llama-2-70Bw3g1283.6571.8829.1Link
Llama-2-70Bw2g644.9669.4420.1Link
Llama-2-70Bw2g1285.2668.7318.9Link
Llama-3-8Bfp166.1468.5813.0-
Llama-3-8Bw4g1286.5068.435.4Link
Llama-3-8Bw3g1287.3466.724.7Link
Llama-3-8Bw2g6412.4758.653.9Link
Llama-3-8Bw2g12813.2558.233.8Link
Llama-3-70Bfp162.8575.33137.8-
Llama-3-70Bw4g1283.1874.5038.9Link
Llama-3-70Bw3g1284.8871.9032.2Link
Llama-3-70Bw2g6413.7566.7023.2Link
Llama-3-70Bw2g12816.7965.0622.0Link
Llama-3-8B-Instructfp168.2968.4313.0-
Llama-3-8B-Instructw4g1288.7667.805.4Link
Llama-3-8B-Instructw3g1289.8366.544.7Link
Llama-3-8B-Instructw2g6416.7758.623.9Link
Llama-3-8B-Instructw2g12818.0257.193.8Link
Llama-3-70B-Instructfp165.3373.78137.8-
Llama-3-70B-Instructw4g1285.7773.5238.9Link
Llama-3-70B-Instructw3g1287.2569.8032.2Link
Llama-3-70B-Instructw2g6412.4865.6023.2Link
Llama-3-70B-Instructw2g12813.4861.7522.0Link

Usage

Please refer https://github.com/OpenGVLab/EfficientQAT for details. These checkpoints can be used to following E2E-AP, as well as be inferenced directly.