CoolFace
Modelpublic

tutuchen2000/Qwen-Image-Edit-Rapid-AIO-NSFW-v23-NVFP4-GGUF

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
1likes539downloads
Model Card

Qwen-Image-Edit-Rapid-AIO-NSFW-v23 — NVFP4 GGUF

This repository contains NVFP4 (4-bit NVIDIA FP4) quantized GGUF components of the Phr00t/Qwen-Image-Edit-Rapid-AIO-NSFW-v23 merged checkpoint, converted for use with ComfyUI-GGUF.

Components

FileSizeDescription
qwen-v23-diffusion-NVFP4.gguf11.5 GBMMDiT diffusion backbone (~20B params) — NVFP4 quantized
text_encoder/text_encoder-NVFP4.gguf3.84 GBQwen2.5-VL-7B text encoder — NVFP4 quantized
vae/vae.safetensors0.25 GBVAE decoder (BF16, original format)

Hardware Requirements

  • —GPU: NVIDIA Blackwell (RTX 50-series) with CUDA 13.0+
  • —VRAM: ~10 GB for diffusion model + text encoder + VAE (NVFP4 saves ~4 GB vs FP8)
  • —System RAM: 32 GB+ recommended

Usage with ComfyUI

  1. 1.Install ComfyUI-GGUF
  2. 2.Place qwen-v23-diffusion-NVFP4.gguf in ComfyUI/models/checkpoints/
  3. 3.Place text_encoder/text_encoder-NVFP4.gguf in ComfyUI/models/clip/
  4. 4.Place vae/vae.safetensors in ComfyUI/models/vae/
  5. 5.Load the checkpoint normally — ComfyUI-GGUF will automatically detect the NVFP4 format

Quantization Details

PropertyValue
Quantization (Diffusion)NVFP4 (NVIDIA FP4, 4-bit)
Quantization (Text Encoder)NVFP4 (NVIDIA FP4, 4-bit)
GGML Quant Type40 (NVFP4)
GGUF File Type39 (LLAMAFTYPEMOSTLY_NVFP4)
Bits Per Weight (Text Encoder)~4.66
BackendCUDA 13.2, compute 12.0
llama.cpp Commitdd7cad7
Source PrecisionF8_E4M3 → F16 → NVFP4

Conversion Process

The original FP8 merged checkpoint was:

  1. 1.Split into three components (diffusion, text encoder, VAE)
  2. 2.Each component's FP8 tensors manually converted to F16 via streaming binary I/O with row-chunked FP8→F16 conversion to avoid OOM on large tensors
  3. 3.Written as F16 GGUF using gguf Python library (v0.19.0)
  4. 4.Quantized to NVFP4 using llama-quantize.exe (commit dd7cad7, CUDA 13.2)
  5. 5.Architecture binary-patched in output GGUF to match expected ComfyUI-GGUF target

Text Encoder Notes

  • —The text encoder (Qwen2.5-VL-7B, 7.75 GB F8_E4M3 → 3.84 GB NVFP4) uses the qwen2 GGUF architecture
  • —Input GGUF was written with both qwen2.* and llama.* metadata keys for quantizer compatibility, then patched to qwen2 after quantization
  • —Row-chunked FP8→F16 conversion (256 rows at a time) prevents OOM on the 545M-element embedding table
  • —Embedding weights are tied (lm_head shares weights with token_embd) — standard for Qwen2 models

Available Quantizations

Currently only NVFP4 is provided. Other quantizations (Q4KM, Q5KM, Q6K, Q80) are possible using the same toolchain.

Notes

  • —This is a NSFW checkpoint. Use at your own discretion.
  • —The text encoder is now also NVFP4 GGUF, saving ~4 GB (was 7.75 GB as safetensors) while remaining compatible with ComfyUI-GGUF.
  • —The VAE remains in safetensors format (BF16) for maximum compatibility.
  • —NVFP4 is a lossy quantization — expect minor quality degradation compared to the original FP8 model.
  • —Requires NVIDIA Blackwell GPU (RTX 50-series) with CUDA 13.0+ for NVFP4 compute support.

Links