weili-0234/Qwen3.5-9B-NVFP4-GPTQ
0284
Qwen3.5-9B-NVFP4-GPTQ
GPTQ (Hessian-corrected PTQ) NVFP4 quantization of Qwen/Qwen3.5-9B in vLLM compressed-tensors nvfp4-pack-quantized format, with calibrated input global scales (W4A4-servable, default load, SM100+). Strong-PTQ baseline of the Qwen3.5-9B standardized campaign; 9B companion of Qwen3.5-27B-NVFP4-GPTQ, produced by the identical recipe.
How this checkpoint was produced — exact reproduction
Inference
vllm serve weili-0234/Qwen3.5-9B-NVFP4-GPTQ --max-model-len 33024