mconcat/Qwopus3.6-27B-v2-NVFP4
Qwopus3.6-27B-v2-NVFP4
Mixed-precision (NVFP4 + FP8 + BF16) quantization of Jackrong/Qwopus3.6-27B-v2, a Claude Opus reasoning-distilled fine-tune of Qwen 3.6 27B.
The hybrid DeltaNet + softmax attention architecture is preserved, the 1-layer MTP head is included as a BF16 sidecar for speculative decoding, and the multimodal processor metadata is kept intact.
Quick start
Requires vLLM ≥ 0.21.0 and a Blackwell-class GPU (SM 10.0+) for native NVFP4 W4A4 inference:
vllm serve mconcat/Qwopus3.6-27B-v2-NVFP4 \
--tensor-parallel-size 1 \
--max-model-len 16384 \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}' \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--trust-remote-codeBenchmarks
Evaluated with lm-evaluation-harness on a single NVIDIA B300 SXM6, 100 samples per task, 0-shot CoT, max_gen_toks=4096:
Accuracy is preserved versus the BF16 source — the GSM8K score is identical to the source and the other tasks match within standard error.
Throughput
Measured on a single NVIDIA B300 SXM6 with vLLM 0.21.0 and torch.compile enabled:
Self-test of tool calling with --tool-call-parser qwen3_coder: passes (model emits well-formed <tool_call>...</tool_call> syntax that the parser extracts correctly).
Quantization
Calibration data: 1024 self-generated reasoning traces from the BF16 source model (256 prompts × 4 generations) spanning math, code, logic, analysis, creative writing, general knowledge, tool calling, and Korean. Generated at temperature=1.0, top_p=0.95.
Files
Total checkpoint size: ~26 GB (down from ~54 GB BF16 source).
License
Apache 2.0 (inherited from the base model).
