JongYeop/Qwen2.5-VL-7B-Instruct-FP4-W4A4-LM-Only
019
Qwen2.5-VL-7B-Instruct-FP4-W4A4-LM-Only
This is a NVFP4 W4A4 quantized version of Qwen/Qwen2.5-VL-7B-Instruct, created using llm-compressor.
Only the LLM decoder is quantized. The Vision Transformer (ViT) encoder remains in BF16 precision.
Model Summary
Quantization Details
- Weights: FP4 (4-bit float), per-tensor-group (group_size=16) with global + local scales, static quantization
- Activations: FP4 (4-bit float), per-tensor-group (group_size=16), dynamic-local quantization
- Ignored:
lm_head(kept in BF16) and all ViT encoder layers (model.visual.*) - Calibration: 512 samples from CNN/DailyMail, max sequence length 2048
Quantization Recipe
quant_stage:
quant_modifiers:
QuantizationModifier:
ignore: ["lm_head", "re:model.visual.*"]
scheme: "NVFP4"
targets: ["Linear"]Usage
With vLLM
export VLLM_ATTENTION_BACKEND=TORCH_SDPA
vllm serve JongYeop/Qwen2.5-VL-7B-Instruct-FP4-W4A4-LM-Only \
--trust-remote-code \
--max-model-len 4096 \
--enforce-eagerWith Transformers
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
"JongYeop/Qwen2.5-VL-7B-Instruct-FP4-W4A4-LM-Only",
torch_dtype="auto",
device_map="auto",
)
processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-7B-Instruct")
messages = [{"role": "user", "content": [
{"type": "image", "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg"},
{"type": "text", "text": "Describe this image in detail."},
]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs, videos=video_inputs, return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=256)
result = processor.batch_decode(output[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(result[0])Environment
Acknowledgments
- Base model by Qwen Team
- Quantization powered by llm-compressor and compressed-tensors
