ISTA-DASLab/Llama-3.1-8B-Instruct-MR-GPTQ-nvfp
027
Llama-3.1-8B-Instruct-MR-GPTQ-nvfp
Model Overview
This model was obtained by quantizing the weights of Llama-3.1-8B-Instruct to NVFP4 data type. This optimization reduces the number of bits per parameter from 16 to 4.5, reducing the disk size and GPU memory requirements by approximately 72%.
Usage
MR-GPTQ quantized models with QuTLASS kernels are supported in the following integrations:
transformerswith these features:- Available in
main(Documentation). - RTN on-the-fly quantization.
- Pseudo-quantization QAT.
vLLMwith these features:- Available in this PR.
- Compatible with real quantization models from
FP-Quantand thetransformersintegration.
Evaluation
This model was evaluated on a subset of OpenLLM v1 benchmarks and Platinum bench. Model outputs were generated with the vLLM engine.
OpenLLM v1 results
Platinum bench results
Below we report recoveries on individual tasks as well as the average recovery.
Recovery by Task
