CoolFace
Modelpublic

ISTA-DASLab/Llama-3.1-8B-Instruct-FPQuant-QAT-MXFP4

sourceHugging Faceupdated 11mo agoView on Hugging Face
0likes156downloads
Model Card

This is the official QAT FP-Quant checkpoint of meta-llama/Llama-3.1-8B-Instruct, produced as described in the **"Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization"** paper.

This model can be run on Blackwell-generation NVIDIA GPUs via QuTLASS and FP-Quant in either transformers or vLLM.

The approximate recipe for training this model (up to local batch size and LR) is available here.

This checkpoint has the following performance relative to the original model and the RTN quantization:

ModelMMLUGSM8kHellaswagWinograndeAvg
meta-llama/Llama-3.1-8B-Instruct72.885.180.077.978.9
RTN62.671.875.272.370.5
QAT (THIS)67.680.378.374.975.3