murilonwt/Qwen2.5-3B-Instruct-NVFP4
032
Qwen2.5-3B-Instruct NVFP4A16 — vLLM / RTX 5060 Ti
This repository contains a locally quantized Qwen2.5-3B-Instruct model using NVFP4A16 / compressed-tensors for vLLM inference on NVIDIA Blackwell GPUs.
All Linear layers (except lm_head) were quantized to NVFP4A16 using llmcompressor + compressed-tensors.
This model card documents the local quantization and test performed on RTX 5060 Ti 16GB.
Quantization Summary
Tested Hardware
Suggested vLLM Command
vllm serve /models/Qwen2.5-3B-Instruct-NVFP4 \
--trust-remote-code \
--served-model-name Qwen2.5-3B \
--max-model-len 32768 \
--gpu-memory-utilization 0.93 \
--max-num-batched-tokens 8192 \
--max-num-seqs 4 \
--tensor-parallel-size 1 \
--enforce-eager \
--port 8000Status
- [x] Quantized to NVFP4A16 (compressed-tensors)
- [x] Tested on RTX 5060 Ti 16 GB
- [x] Tested with vLLM v0.22.0 Docker image
- [x] Loads successfully in vLLM
