CoolFace
Modelpublic

YihongJin/Qwen3-Omni-30B-A3B-Instruct-NVFP4-W4A4-full-thinker-awqclip

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
1likes6.3kdownloads
Model Card

Qwen3-Omni-30B-A3B-Instruct — NVFP4 W4A4 (full thinker, awq_clip text calibration)

ModelOpt NVFP4 W4A4 quantization of Qwen/Qwen3-Omni-30B-A3B-Instruct.

The entire thinker text body — attention QKV/O + MoE experts — is quantized to NVFP4 with FP8 per-tensor input scales. Embeddings, norms, the MoE router (mlp.gate), lm_head, the audio encoder, the vision encoder, the talker, and code2wav stay in BF16.

Total checkpoint size27.6 GiB (vs 66 GiB BF16)
Size reduction~58%
Hardware requirementNVIDIA Blackwell (sm_100+) for native FP4 GEMM
Kernel pathFlashInfer Cutlass NvFp4 Linear + FlashInfer TRT-LLM NvFp4 MoE
Exported safetensors NaN bytes0 (calibrated with the ModelOpt-side fix; see Mitigations below)

Accuracy — Daily-Omni (full split, n=1197)

Benchmarked on NVIDIA B200 (HBM3e, 183 GiB) via the in-tree `run_qwen_omni_acc_benchmark.py` harness. BF16 baseline run on the same hardware, same harness, same prompt set. CUDA graphs enabled (no --enforce-eager).

ModelDaily-Omni overalln correctΔ vs BF16
Qwen/Qwen3-Omni-30B-A3B-Instruct (BF16)0.694831 / 1197—
This checkpoint (W4A4 NVFP4)0.675808 / 1197-1.9 pp

Accuracy — OmniBench (n=200, via evalscope)

ModelOmniBench mean_accΔ vs BF16
BF160.455—
W4A4 NVFP40.450-0.5 pp

Throughput / latency — vLLM serve on B200

Synthetic random workload (--random-input-len 256, --random-output-len 512, --num-prompts 200, no --enforce-eager, CUDA graphs captured for prefill + decode batch sizes 1..256).

ConcurrencyMetricBF16**W4A4 NVFP4**W4A4 vs BF16
1Output tok/s156150-4%
1TPOT (ms)5.875.84tied
1TTFT (ms)72111+54% (BF16 wins prefill)
8Output tok/s8691002+15%
8TPOT (ms)8.336.73-19%
32Output tok/s20832766+33%
32TPOT (ms)12.769.12-29%
64Output tok/s31974112+29%
64TPOT (ms)15.8710.96-31%
128Output tok/s49235665+15%
128TPOT (ms)18.0115.36-15%

W4A4 wins on every metric at conc=8+. The single-stream TTFT loss reflects that FP4 prefill has small dequant overhead that BF16 native tensor cores don't pay; once concurrency saturates the GEMM, FP4 bandwidth savings dominate and W4A4 pulls ahead by 15-33% in tokens/s and 15-31% in per-token latency.

Stability

30-minute stress at concurrency=32: 3200 successful requests, 0 errors, 0 NaN collapses (regression guard on the !!!! failure mode that motivated vllm-omni#4025).

Calibration recipe

  • —Base model: Qwen/Qwen3-Omni-30B-A3B-Instruct in bfloat16.
  • —ModelOpt: nvidia-modelopt==0.44.0 with the ModelOpt-side calibration fix applied (see Mitigations below). Exported safetensors contain ModelOpt's raw output with no Python post-processing (--skip-post-export-nan-clamp).
  • —Quant config: mtq.NVFP4_DEFAULT_CFG + algorithm awq_clip.
  • —Calibration set: 1024 prompts from `HuggingFaceH4/ultrachat_200k` train_sft, chat-templated through the Qwen3-Omni tokenizer, truncated to 512 tokens each. Text-only — multimodal samples were not used for this checkpoint.
  • —Excluded patterns: *audio_tower*, *visual*, *talker*, *code2wav*, *lm_head*, *mlp.gate* (router stays BF16 to avoid expert-routing drift).
  • —Calibration time: ~2.4 h on a single RTX PRO 6000 Blackwell WS.

Inference

python
from vllm_omni import Omni
omni = Omni(model="YihongJin/Qwen3-Omni-30B-A3B-Instruct-NVFP4-W4A4-full-thinker-awqclip")

OpenAI-compatible server (recommended config for max throughput):

bash
vllm serve YihongJin/Qwen3-Omni-30B-A3B-Instruct-NVFP4-W4A4-full-thinker-awqclip \
    --omni --port 8000
Do not pass --enforce-eager for benchmarks. CUDA graphs amortize kernel launch overhead and unlock the FP4 throughput wins above; with --enforce-eager set, W4A4 TPOT degrades 10x relative to the CUDA-graph configuration.

Compute requirement: sm_100+ (Blackwell — B100, B200, RTX 5090, RTX Pro 6000) for the native FlashInfer FP4 GEMM kernel.

ModelOpt 0.44 NaN regression — two mitigation paths

ModelOpt 0.44's float32 -> torch.float8_e4m3fn cast of per-block weight_scale occasionally emits literal NaN bytes (E4M3 encoding 0x7F / 0xFF) when the pre-cast scale rounds above the FP8 max of 448 after the global-scale division. A single NaN byte in any weight_scale propagates through the FlashInfer FP4 GEMM into the residual stream and collapses the served model output to !!!!. Two complementary fixes:

  1. 1.Calibration-time (ModelOpt-side): clamp the pre-cast values to torch.finfo(torch.float8_e4m3fn).max before every .to(torch.float8_e4m3fn) at the two cast sites in modelopt/torch/quantization/qtensor/nvfp4_tensor.py and modelopt/torch/export/quant_utils.py. This checkpoint was calibrated with that ModelOpt 0.44 patch applied — exported safetensors contain 0 NaN bytes. An upstream PR to NVIDIA/TensorRT-Model-Optimizer is in progress.
  1. 1.Load-time (vllm-omni-side): vllm-project/vllm-omni#4025 installs a defensive override of ModelOptNvFp4LinearMethod.process_weights_after_loading that scans weight_scale for NaN bytes and clamps them to FP8 E4M3 max at worker init. Because this checkpoint is already clean, the override is a no-op safety net here; it primarily protects other in-the-wild W4A4 NVFP4 checkpoints that were exported with vanilla ModelOpt 0.44 and currently serve as !!!!. Self-extinguishes once vllm-omni's vllm pin includes the corresponding upstream vLLM fix; can be disabled with VLLM_OMNI_SKIP_NVFP4_NAN_CLAMP=1 for diagnostics.

License

Apache-2.0 (inherits from the base Qwen3-Omni-30B-A3B-Instruct model).