iapp/openthai2.0-qwen3.8-27b-NVFP4
23.6k
v2.0.2 NVFP4
v2.0.2 NVFP4
Restore MTP draft head (was missing: acceptance 1.00, 1.6x slower with speculative decoding). Adds model-mtp.safetensors, maps it in the index, ignores re:.*mtp.* in quantization_config. Verified: mean acceptance length 1.68 under vLLM qwen3_5_mtp.
v2.0.1 NVFP4
Re-quantized from release winner
Quant card
NVFP4 (llmcompressor, ultrachat-calibrated; vision+MTP bf16)
initial commit
