Abiray/Qwen3.6-27B-NVFP4
12125
Qwen3.6-27B-NVFP4
NVFP4 quantized version of Qwen/Qwen3.6-27B by Abiray using custom Blackwell NVFP4 GEMM kernels
55.6 GB → 19.7 GB (0.35x) with vision tower preserved in BF16.
NVFP4 Quantization Details
Recipe
QuantizationModifier:
targets: [Linear]
ignore: [lm_head, 're:.*visual.*', 're:.*mlp.gate$', 're:.*mlp.shared_expert_gate$']
scheme: NVFP4