CoolFace
Modelpublic

lite-infer/flux.1-schnell-nunchaku-lite-nvfp4_r32-bnb4-text-encoder

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes11downloads
Model Card

FLUX.1 Schnell Nunchaku Lite NVFP4 r32

Diffusers-loadable conversion of:

  • —Base model: black-forest-labs/FLUX.1-schnell
  • —Source repo: nunchaku-ai/nunchaku-flux.1-schnell
  • —Source checkpoint: svdq-fp4_r32-flux.1-schnell.safetensors

The transformer uses quant_method: nunchaku_lite, NVFP4 SVDQ with group size 16, runtime rank 64, 418 SVDQ targets, and 76 AWQ W4A16 targets. The CLIP encoder is copied from the base model and T5 text_encoder_2 is BitsAndBytes 4-bit NF4. Fused QKV modules are split in logical tensor layout; single-block proj_out is merged from attention and MLP projections; low-rank tensors are logically padded to rank 64. NVFP4 outer scales are reconciled without overflowing FP8 group scales.

Benchmark

CheckpointLatencyMax VRAM
Converted Diffusers Nunchaku Lite NVFP4 r32 + BNB4 T51.34 s (stdev 0.03 s)16.73 GiB
Original Nunchaku NVFP4 r32 + BF16 T50.81 s (stdev 0.00 s)20.99 GiB

RTX 5090, 1024×1024, 4 steps, guidance scale 0.0, one warmup and three measured runs, full GPU placement. VRAM is peak total device usage sampled with nvidia-smi, including allocations outside PyTorch's caching allocator. The native row uses the base model's BF16 T5 encoder.

Output Comparison

[image]

Both images use the same prompt, seed 0, scheduler, resolution, and step count. The native and converted NVFP4 images use identical inputs.

Run

Requires the Hugging Face kernels package and a Blackwell NVIDIA GPU for NVFP4 kernels.

python
import torch
from diffusers import FluxPipeline

pipe = FluxPipeline.from_pretrained(
    "lite-infer/flux.1-schnell-nunchaku-lite-nvfp4_r32-bnb4-text-encoder",
    torch_dtype=torch.bfloat16,
).to("cuda")

image = pipe(
    prompt='A cinematic photograph of a red fox standing in a misty forest at sunrise, detailed fur, volumetric light',
    generator=torch.Generator("cuda").manual_seed(0),
    width=1024,
    height=1024,
    num_inference_steps=4,
    guidance_scale=0.0,
).images[0]
image.save("output.png")