CoolFace
Modelpublic

RESMP-DEV/Qwen3-Next-80B-A3B-Thinking-NVFP4

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
10likes99downloads
Model Card

Qwen3-Next-80B-A3B-Thinking-NVFP4

Quantized version of [Qwen/Qwen3-Next-80B-A3B-Thinking](https://huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking) using LLM Compressor and the NVFP4 (E2M1 + E4M3) format.

This time it actually works! We think

This should be the start of a new series of hopefully optimal NVFP4 quantizations as capable cards continue to grow out in the wild.


Model Summary

PropertyValue
Base modelQwen/Qwen3-Next-80B-A3B-Thinking
QuantizationNVFP4 (FP4 microscaling, block = 16, scale = E4M3)
MethodPost-Training Quantization with LLM Compressor
ToolchainLLM Compressor
Hardware targetNVIDIA Blackwell (Untested on RTX cards) / GB200 Tensor Cores
PrecisionWeights & activations = FP4 • Scales = FP8 (E4M3)
MaintainerRESMP.DEV

Description

This model is a drop-in replacement for Qwen/Qwen3-Next-80B-A3B-Thinking that runs in NVFP4 precision. Accuracy remains within ≈ 1 % of the FP8 baseline on standard reasoning and coding benchmarks.