CoolFace
Modelpublic

stuqiu/nunchaku-qwen-image-2512

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
3likes171downloads
Model Card

nunchaku-qwen-image-2512 (SVDQuant W4A4 NVFP4)

4-bit SVDQuant of Qwen/Qwen-Image-2512 for the nunchaku inference engine. Same file layout and usage as the official nunchaku-tech/nunchaku-qwen-image release — load it with stock from_pretrained, no extra code.

  • —Weights: W4 (NVFP4 e2m1) + low-rank branch (rank 32 or rank 128), activations A4 (NVFP4)
  • —Modulation layers: W4 AWQ (symmetric, group 64)
  • —Requires: Blackwell GPU (RTX 50-series sm120 / GB10 sm121) for NVFP4
  • —Measured: 238 ms/step on RTX 5090, 1.06 s/step on DGX Spark GB10 (see benchmarks below); quality on par with the official base-model release (incl. Chinese text rendering, the point of the 2512 finetune)

Models

Data TypeRankModel NameComment
NVFP4r32`svdq-fp4_r32-qwen-image-2512.safetensors`~12 GB · smaller / faster · default
r128`svdq-fp4_r128-qwen-image-2512.safetensors`~13 GB · higher fidelity (finer detail & text edges)

Both checkpoints are W4A4 NVFP4 and differ only in the SVDQuant low-rank branch size. NVFP4 needs a Blackwell GPU (RTX 50-series sm120 / GB10 sm121); INT4 builds are not provided for this finetune yet.

Usage

Identical to the official nunchaku Qwen-Image example — two lines changed:

python
import torch
from diffusers import QwenImagePipeline
from nunchaku.models.transformers.transformer_qwenimage import NunchakuQwenImageTransformer2DModel

from nunchaku.utils import get_precision

rank = 32  # 32 (smaller/faster) or 128 (higher fidelity)
transformer = NunchakuQwenImageTransformer2DModel.from_pretrained(
    f"stuqiu/nunchaku-qwen-image-2512/svdq-{get_precision()}_r{rank}-qwen-image-2512.safetensors"
)
pipe = QwenImagePipeline.from_pretrained(
    "Qwen/Qwen-Image-2512", transformer=transformer, torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
    prompt="夜晚的中式茶馆门口,门匾上写着「天下为公」四个金色大字",
    num_inference_steps=50, true_cfg_scale=4.0,
).images[0]
image.save("out.png")

Benchmarks

1024×1024, 50 steps, stock pipeline:

HardwarePrecisionper-step50-step imageGPU mem after load
RTX 5090 (sm_120)this model (W4A4)238 ms (~5.0 it/s)11.9 s26.8 GiB
RTX 5090 (sm_120)BF16n/a — model does not fit in 32 GB, requires CPU offloading (much slower)
DGX Spark GB10 (sm_121)this model (W4A4)1.06 s~55 s—
DGX Spark GB10 (sm_121)BF161.80 s~90 s—

Pipeline load 4.5 s, warmup step 2.6 s (RTX 5090).

Samples

Generated on RTX 5090 with the rank-32 checkpoint (50 steps, 1024×1024); see the rank 32 vs rank 128 comparison below:

[image][image]
8k high quality portrait of a beautiful woman, sharp focus, detailed face8k photorealistic landscape, golden hour, mountains and lake

[image]

夜晚的中式茶馆门口,门匾上写着「天下为公」四个金色大字

rank 32 vs rank 128 (seed 42, 50 steps, generated on DGX Spark GB10)

rank 32rank 128
[image][image]

夜晚的中式茶馆门口,门匾上写着「天下为公」四个金色大字

rank 32rank 128
[image][image]

A cat sitting on a chair

rank 32rank 128
[image][image]

A 20-year-old East Asian girl with delicate features and large bright brown eyes, naturally wavy long hair, fair skin and light makeup, in a cute bright-colored dress, standing indoors at an anime convention surrounded by banners and posters — casual iPhone-snapshot look.

Quantization pipeline

deepcompressor SVDQuant calibration (64 samples, smoothing grid 10, rank-32 low-rank branches) for attention/MLP, official symmetric AWQ recipe (scale = absmax/7, zero ≡ 7) for modulation, plus the super-channel fold above.

Acknowledgements