CoolFace
Modelpublic

tonera/Juggernaut-XL_v9_RunDiffusionPhoto_v2

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
1likes99downloads
Model Card

Model Card (SVDQuant)

Language: English | 中文

Model Name

  • Model repo: tonera/Juggernaut-XL_v9_RunDiffusionPhoto_v2
  • Base (Diffusers weights path): tonera/Juggernaut-XL_v9_RunDiffusionPhoto_v2 (repo root)
  • Quantized UNet weights: tonera/Juggernaut-XL_v9_RunDiffusionPhoto_v2/svdq-<precision>_r32-Juggernaut-XL_v9_RunDiffusionPhoto_v2.safetensors

Quantization / Inference Tech

  • Inference engine: Nunchaku (https://github.com/nunchaku-ai/nunchaku)

Nunchaku is a high-performance inference engine for 4-bit (FP4/INT4) low-bit neural networks. Its goal is to significantly reduce VRAM usage and improve inference speed while preserving generation quality as much as possible. It implements and productionizes post-training quantization methods such as SVDQuant, and reduces the overhead introduced by low-rank branches via operator/kernel fusion and other optimizations.

The SDXL quantized weights in this repository (e.g. svdq-*_r32-*.safetensors) are intended to be used with Nunchaku for efficient inference on supported GPUs.

Quantization Quality (fp8)

text
PSNR: mean=21.6242 p50=22.2476 p90=24.7405 best=27.4656 worst=13.4446 (N=25)
SSIM: mean=0.776742 p50=0.800346 p90=0.882404 best=0.896764 worst=0.401949 (N=25)
LPIPS: mean=0.234511 p50=0.21222 p90=0.326609 best=0.0975928 worst=0.634302 (N=25)

Performance

Below is the inference performance comparison (Diffusers vs Nunchaku-UNet).

  • Inference config: bf16 / steps=30 / guidance_scale=5.0
  • Resolutions (5 images each, batch=5): 1024x1024, 1024x768, 768x1024, 832x1216, 1216x832
  • Software versions: torch 2.9 / cuda 12.8 / nunchaku 1.1.0+torch2.9 / diffusers 0.37.0.dev0
  • Optimization switches: no torch.compile, no explicit cudnn tuning flags

Cold-start performance (end-to-end for the first image)

GPUMetricDiffusersNunchakuSpeedupGain
RTX 5090load3.505s3.432s1.02x+2.1%
RTX 5090cold_infer2.944s2.447s1.20x+16.9%
RTX 5090cold_e2e6.449s5.880s1.10x+8.8%
RTX 3090load3.787s3.442s1.10x+9.1%
RTX 3090cold_infer7.503s5.231s1.43x+30.3%
RTX 3090cold_e2e11.290s8.673s1.30x+23.2%

Steady-state performance (5 consecutive images after warmup)

GPUMetricDiffusersNunchakuSpeedupGain
RTX 5090total (5 images)12.937s9.813s1.32x+24.2%
RTX 5090avg (per image)2.587s1.963s1.32x+24.2%
RTX 3090total (5 images)33.413s22.975s1.45x+31.2%
RTX 3090avg (per image)6.683s4.595s1.45x+31.2%

Notes:

  • The longer load time on RTX 3090 is due to extra one-time processing when loading quantized weights.
  • During inference (cold_infer and steady-state), Nunchaku shows clear speedups on both GPUs.

Nunchaku Installation Required

  • Official installation docs (recommended source of truth): https://nunchaku.tech/docs/nunchaku/installation/installation.html

(Recommended) Install the official prebuilt wheel

  • Prerequisite: PyTorch >= 2.5 (follow the wheel requirements)
  • Install Nunchaku wheel: choose a wheel matching your torch/cuda/python versions from GitHub Releases / HuggingFace / ModelScope (note cp311 means Python 3.11):
  • https://github.com/nunchaku-ai/nunchaku/releases
bash
# Example (select the correct wheel URL for your torch/cuda/python versions)
pip install https://github.com/nunchaku-ai/nunchaku/releases/download/vX.Y.Z/nunchaku-X.Y.Z+torch2.9-cp311-cp311-linux_x86_64.whl
  • Tip (RTX 50 series): typically prefer CUDA >= 12.8, and prefer FP4 models for compatibility/performance (follow official docs).

Usage Example (Diffusers + Nunchaku UNet)

python
import torch
from diffusers import StableDiffusionXLPipeline

from nunchaku.models.unets.unet_sdxl import NunchakuSDXLUNet2DConditionModel
from nunchaku.utils import get_precision

MODEL = "Juggernaut-XL_v9_RunDiffusionPhoto_v2"  # Replace with the actual model name before publishing (e.g. zavychromaxl_v100)
REPO_ID = f"tonera/{MODEL}"

if __name__ == "__main__":
    unet = NunchakuSDXLUNet2DConditionModel.from_pretrained(
        f"{REPO_ID}/svdq-{get_precision()}_r32-{MODEL}.safetensors"
    )

    pipe = StableDiffusionXLPipeline.from_pretrained(
        f"{REPO_ID}",
        unet=unet,
        torch_dtype=torch.bfloat16,
        use_safetensors=True,
    ).to("cuda")

    prompt = "Make Pikachu hold a sign that says 'Nunchaku is awesome', yarn art style, detailed, vibrant colors"
    image = pipe(prompt=prompt, guidance_scale=5.0, num_inference_steps=30).images[0]
    image.save("sdxl.png")