smoothquant
smoothquant-kivi-w8a8kv8
SmoothQuant W8A8 + KIVI-INT8 KV (Llama-3.1-8B-Instruct)
w8_of_w8a8_smoothquant_llama_31_8b/ — W8 weights
The INT8 weights of the SmoothQuant W8A8 model (per-channel symmetric; weight = int8 × scale).
Stored per layer: layer_0.safetensors … layer_31.safetensors + embeddings.safetensors.
The 7 linears per layer are quantized (int8 weight + fp16 per-output-channel scale); everything else stays fp16:
key
dtype
shape
self_attn.q_proj.weight… See the full description on the dataset page: https://huggingface.co/datasets/jsyeom/smoothquant-kivi-w8a8kv8.smoothquant-dataset
