modal-labs/Qwen3.8-2.4T-A95B-NVFP4
1401
Qwen3.8-2.4T-A95B-NVFP4
NVFP4 quantization of Qwen/Qwen3.8-2.4T-A95B. Routed experts in NVFP4 (group size 16): attention, linear attention, shared experts, router gates, embeddings, and lm_head keep their original precision. Text-only, 262144 native context extended to 1M using RoPE at serving time.
