CoolFace
Modelpublic

JongYeop/Qwen2.5-VL-3B-Instruct-FP4-W4A4-LM-Only

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes21downloads
Model Card

Qwen2.5-VL-3B-Instruct-FP4-W4A4-LM-Only

This is a NVFP4 W4A4 quantized version of Qwen/Qwen2.5-VL-3B-Instruct, created using llm-compressor.

Only the LLM decoder is quantized. The Vision Transformer (ViT) encoder remains in BF16 precision.

Model Summary

PropertyValue
Base ModelQwen/Qwen2.5-VL-3B-Instruct
QuantizationNVFP4 W4A4 (4-bit float weights, 4-bit float activations)
Quantization ScopeLLM decoder only (ViT encoder in BF16)
StrategyPer-tensor-group (group_size=16), symmetric (minmax observer)
Formatcompressed-tensors (nvfp4-pack-quantized)
Model Size~3.9 GB (1 shard)
Ignored Layerslm_head, all model.visual.* layers
Toolllm-compressor v0.7.1
Supported RuntimevLLM (with compressed-tensors)

Quantization Details

  • —Weights: FP4 (4-bit float), per-tensor-group (group_size=16) with global + local scales, static quantization
  • —Activations: FP4 (4-bit float), per-tensor-group (group_size=16), dynamic-local quantization
  • —Ignored: lm_head (kept in BF16) and all ViT encoder layers (model.visual.*)
  • —Calibration: Data-free (no calibration dataset required)

Quantization Recipe

yaml
quant_stage:
  quant_modifiers:
    QuantizationModifier:
      ignore: ["lm_head", "re:model.visual.*"]
      scheme: "NVFP4"
      targets: ["Linear"]

Usage

With vLLM

bash
export VLLM_ATTENTION_BACKEND=TORCH_SDPA

vllm serve JongYeop/Qwen2.5-VL-3B-Instruct-FP4-W4A4-LM-Only \
    --trust-remote-code \
    --max-model-len 4096 \
    --enforce-eager

With Transformers

python
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info

model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    "JongYeop/Qwen2.5-VL-3B-Instruct-FP4-W4A4-LM-Only",
    torch_dtype="auto",
    device_map="auto",
)
processor = AutoProcessor.from_pretrained("Qwen/Qwen2.5-VL-3B-Instruct")

messages = [{"role": "user", "content": [
    {"type": "image", "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg"},
    {"type": "text", "text": "Describe this image in detail."},
]}]

text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs, videos=video_inputs, return_tensors="pt").to(model.device)

output = model.generate(**inputs, max_new_tokens=256)
result = processor.batch_decode(output[:, inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(result[0])

Environment

ComponentVersion
llm-compressor0.7.1
compressed-tensors0.11.0
transformers4.55.2
torch2.8.0+cu128
GPUNVIDIA RTX PRO 6000 (98GB)

Acknowledgments