CoolFace
Modelpublic

nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-MTP-NVFP4

sourceHugging Faceapache-2.0updated 16d agoView on Hugging Face
0likes
Model Card

<!-- NVFP4CARDV2ENSTART -->

Qwen3.8-27B EfficientThink — NVFP4 + BF16 MTP

[image]

BF16 / FP8 main model · GGUF variants

<!-- DFLASHSTATICFP8V1EN_START -->

True static FP8 DFlash2 draft

The bundled DFlash2 draft is now a pre-quantized static FP8 compressed-tensors checkpoint, not a BF16 checkpoint carrying an FP8 directory label. model.safetensors is 2,407,027,720 bytes (SHA256 1f3636a32d866f8ebc7f422d63f9247126ebb6d2566d3e0da327d81dd8fa25d1). Its audited tensor set contains 20 FP8 E4M3 weights with 20 FP32 scales and 61 retained BF16 tensors.

Load it explicitly with:

text
--speculative-draft-model-quantization compressed-tensors

Matched DGX Spark checks used the same W4A4 target, 15 prompts, XH, 256 generated tokens, and 8 draft tokens. All 15 requests completed in every cell.

DraftConcurrencyAggregate tok/sDFlash acceptanceMean accepted / 8Request errors
BF16 referenceC129.7538.77%3.720
Static FP8C130.2635.86%3.510
BF16 referenceC468.5032.71%3.290
Static FP8C475.9433.55%3.350

C1 acceptance is the mean of same-run log snapshots; C4 acceptance is the post-run SGLang metrics gauge. The C4 static draft improved aggregate throughput by about 10.9% over BF16 in this matched check. This short fixed-length test validates serving behavior; it does not replace the formal capability scores elsewhere in this card. <!-- DFLASHSTATICFP8V1EN_END -->

Research disclaimer: This experimental release is provided solely to study the technical feasibility and behavioral effects of refusal-tendency dissolution. It is not a comprehensive safety conclusion, an endorsement of unrestricted use, or professional advice. Users are responsible for lawful and appropriate use and for independently verifying model outputs.

The figure presents this model's W4A4, W4A4+W8A8, W4A16, and W8A16 variants. Fast/Mixed use C24 and W4A16 uses C16; W8A16 uses C16 for GPQA/LCB and C24 for MMLU. All 12 / 11 GPQA timeouts remain in the Fast/Mixed 198-question failure denominators; no anomalous question was dropped.

<!-- INT8W8A8QATHFXETV21 --> <!-- INT8W8A8QATV2EN_START -->

True-QAT INT8 W8A8 | dynamic INT8 activations

[image]

Download directory: `INT8-W8A8-QAT/`. Matching main/standalone repository on this platform.

Files and precision

ComponentPathPrecision / roleSize
Main modelINT8-W8A8-QAT/model-00001-of-00008.safetensors … model-00008-of-00008.safetensorsTrue-QAT INT8 W8A8; dynamic INT8 activations29.48 GB
Vision + native MTPINT8-W8A8-QAT/vision-mtp-bf16.safetensors333 BF16 vision tensors + 15 BF16 MTP tensors, with 348 real index mappings1.77 GB
Complete inventoryINT8-W8A8-QAT/manifest.json and INT8-W8A8-QAT/SHA256SUMSRoles, bytes, and SHA256 for the current 29-file directory—
Structured evaluation`INT8-W8A8-QAT/evaluation/formal-quality-and-performance.json`Formal scores, reasoning statistics, and the complete research record—

Training and export method

  • —64-layer Qwen3.8-27B multimodal architecture with 1,599 entries in the published model index.
  • —3,200 QAT optimizer steps; all 400 language linear tensors recorded non-zero gradients and are published with INT8 weights.
  • —Dynamic INT8 activations; 247 items entered the accepted training set.
  • —BF16 scales were losslessly exported as F32; the vision tower and native MTP remain BF16.

Formal capability and reasoning results

Protocol: 1× RTX PRO 6000 Blackwell 96GB, vLLM 0.28.0 + native MTP3, C20, BF16 KV, reasoning_effort=xhigh, a 32,768-token output cap, and a 1,800-second request timeout. The long-output formal suite uses C20 because C24 did not leave enough KV capacity for the full suite.

SuiteScoreMean reasoningP50 / P90>8K / >16K32K trunc.Empty final / unparseable
GPQA162/198 (81.82%)9,5304,699.5 / 32,76773 / 412323 / 23
MMLU451/500 (90.20%)837203 / 1,906.19 / 300 / 0
LCB73/100 (73.00%)13,9218,347 / 32,76850 / 412323 / 23

Request / HTTP / capture / grader errors are all 0. LCB has 0 code timeouts and 0 syntax errors, plus 1 runtime error. IPC-v4 uniformly regraded the original 100 answers without issuing new model requests.

24 short-output cells for bundled runtime paths

Protocol: 1,024 input / 256 output, warmup plus 3 trials. This measures short fixed-length serving throughput, not long-reasoning speed. All 24 bare/MTP3 cells completed with 0 request errors.

  • —Highest measured throughput for this tier: vLLM MTP3 C24 at 662 tok/s, 56.28% acceptance, about 27.6 tok/s/request.
  • —SGLang MTP3 C24: 654 tok/s at 55.17% acceptance.
  • —The release bundles and recommends native MTP3 only; the structured evaluation file preserves the complete historical research record.
Framework / modeCAggregate tok/sPer-request tok/sAcceptanceTTFT P50Latency P50Peak GPUErrors
vLLM bareC13231.9—0.16s8.03s85.8 GiB0
vLLM bareC411328.2—0.56s9.04s86.1 GiB0
vLLM bareC821326.6—1.01s9.55s86.1 GiB0
vLLM bareC1636322.7—1.59s11.15s86.1 GiB0
vLLM bareC2042321.2—1.87s11.95s86.1 GiB0
vLLM bareC2447619.8—2.16s12.71s86.1 GiB0
vLLM MTP3C15655.947.48%0.18s4.58s85.8 GiB0
vLLM MTP3C420250.454.87%0.56s4.62s86.0 GiB0
vLLM MTP3C835644.557.08%1.09s5.27s86.0 GiB0
vLLM MTP3C1654434.057.31%1.70s6.88s86.0 GiB0
vLLM MTP3C2061931.056.30%2.00s7.77s86.0 GiB0
vLLM MTP3C2466227.656.28%2.31s8.63s86.0 GiB0
SGLang bareC14545.2—0.14s5.67s87.9 GiB0
SGLang bareC415839.4—0.43s6.49s88.1 GiB0
SGLang bareC828735.9—0.69s7.13s88.1 GiB0
SGLang bareC1646929.3—1.22s8.73s88.1 GiB0
SGLang bareC2053726.8—1.48s9.53s88.1 GiB0
SGLang bareC2459724.9—1.74s10.29s88.1 GiB0
SGLang MTP3C18686.167.86%0.15s2.97s86.5 GiB0
SGLang MTP3C423458.554.03%0.43s3.83s86.7 GiB0
SGLang MTP3C837947.451.95%0.72s4.96s86.7 GiB0
SGLang MTP3C1657836.155.06%1.25s6.56s86.7 GiB0
SGLang MTP3C2061230.654.42%1.52s8.06s86.7 GiB0
SGLang MTP3C2465427.355.17%1.79s8.90s86.7 GiB0

Verified launch paths

bash
cd INT8-W8A8-QAT
bash scripts/serve-vllm-mtp3.sh
ScriptPurpose
scripts/serve-vllm-bare.shvLLM bare
scripts/serve-vllm-mtp3.shvLLM native MTP3; recommended throughput path
scripts/serve-sglang-bare.shSGLang bare
scripts/serve-sglang-mtp3.shSGLang native MTP3

All four bundled paths passed text, image, and real-video smoke on the same model hash. SGLang compressed-tensors INT8 on Blackwell SM120/121 uses the bundled runtime/sglang-sm120-int8-compat/ compatibility layer.

<!-- INT8W8A8QATV2EN_END -->

Formal capability results

VariantGPQA 198MMLU 500LCB 100
W4A4 (NVFP4 Fast)158/198 (79.80%)447/500 (89.40%)74/100 (74.00%)
W4A4+W8A8 (NVFP4 Mixed Precision)168/198 (84.85%)458/500 (91.60%)75/100 (75.00%)
W4A16161/198 (81.31%)457/500 (91.40%)74/100 (74.00%)
W8A16159/198 (80.30%)450/500 (90.00%)78/100 (78.00%)

Fast/Mixed use the same C24 runtime protocol for all three scores. The 12 / 11 GPQA timeouts remain in the complete denominators; anomalous questions were neither removed nor rescored. Results from different runtime formats and decoders are not controlled measurements of quantization loss.

<!-- NVFP4BENEFITSV4ENSTART -->

Mixed precision versus fast: scores and reasoning cost

MetricW4A4 FastW4A4 + W8A8 MixedChange
GPQA score158/198 (79.80%)168/198 (84.85%)+10 correct / +5.05pp
GPQA mean reasoning9,1238,400-7.9%
GPQA P50 / P904,928.5 / 24,893.04,278.0 / 24,402.6—
MMLU score447/500 (89.40%)458/500 (91.60%)+11 correct / +2.20pp
MMLU mean reasoning963848-11.9%
MMLU P50 / P90225.0 / 2,285.5216.5 / 1,668.9—
LCB score74/100 (74.00%)75/100 (75.00%)+1 correct / +1.00pp
LCB mean reasoning14,60213,647-6.5%
LCB P50 / P909,626.5 / 32,769.07,441.5 / 32,769.0—

Under this C24 protocol, mixed precision answers 10 more GPQA and 11 more MMLU questions correctly, while LCB accuracy is unchanged; mean reasoning tokens are lower in all three suites. This compares quantized variants, not the official base against post-training. MMLU/LCB means cover 500/100 questions; percentage changes use unrounded means. <!-- NVFP4BENEFITSV4ENEND -->

<!-- W4A16FINALV1ENSTART -->

W4A16 formal C16 full-suite results

Protocol: 1× RTX PRO 6000 96GB, SGLang + official BF16 MTP, C16, xhigh, a 32,768-token output cap, and a 1,800-second request timeout. This is not a controlled equal-concurrency comparison against the C24 W4A4 runs.

SuiteFinal scoreMean reasoningP50 / P90>8K / >16K32K trunc.Other anomalies
GPQA161/198 (81.31%)11,4046,795.5 / 32,76792 / 572729 empty final-channel outputs; 27 unparseable responses; 0 request errors
MMLU457/500 (91.40%)802225 / 1,5549 / 100 request errors; 0 empty finals
LCB74/100 (74.00%)14,5118,513.5 / 32,76951 / 42260 request errors/timeouts; 26 empty-code cases

<!-- W4A16GPQABOUNDARYV1EN_START -->

GPQA scoring: Final 161/198 (81.31%) across all 198 questions, with 0 request errors and 0 timeouts. Parsing accepts only a non-empty final channel or a complete explicit Final Answer: A/B/C/D on the last non-empty reasoning line.

<!-- W4A16GPQABOUNDARYV1EN_END -->

LCB difficulty: Easy 23/23 (100%), Medium 29/31 (93.55%), Hard 22/46 (47.83%). All 26 length-limited outputs and all 26 empty-code cases remain failures in the 100-question denominator.

All three NVFP4 MMLU results use the same offline final-content-strict-single-letter-v2 rescore: Fast 447/500 (89.40%), Mixed Precision 458/500 (91.60%), and W4A16 457/500 (91.40%); generation outputs are unchanged. The retired 430/449/439 scores are not used.

W4A16 MTP short-output concurrency

The fixed short-output grid tested only C1/C4/C8/C16/C24; C2 and C32 were not tested. Aggregate throughput is rounded to whole tokens/s.

ConcurrencyAggregate tok/sMTP acceptanceAccepted draft / verificationCommitted output / verification
C112772.84%2.1853.160
C445270.20%2.1063.103
C876471.04%2.1313.127
C16 (balanced)1,21674.28%2.2293.228
C24 (max throughput)1,30470.60%2.1183.114

All five cells had 0 request errors, 0 timeouts, and 0 empty outputs. Punctuation-collapse manual review was not part of this short sweep, so no zero claim is made for that field. C16 provides the highest acceptance while retaining 1,216 tok/s; C24 is the highest measured aggregate-throughput point. <!-- W4A16FINALV1ENEND -->

The tested W4A16 SGLang settings are --quantization modelopt_mixed, EAGLE, steps=3, top-k=1, draft tokens=4, BF16 dtype/KV, and a 65,536-token context. The four-mode vLLM smoke still covers only Fast and Mixed Precision.

<!-- W8A16RELEASEV1ENSTART -->

W8A16 quality-oriented FP8 package

The complete W8A16 package is available at W8A16/: 24 files / 38,477,562,434 bytes, including the manifests. Download the entire directory and do not mix it with W4A4/, W4A4+W8A8/, or W4A16/.

Format note: W8A16 uses block-wise FP8 E4M3 weights with BF16 activations and KV cache. It is grouped in this repository family for distribution, but it is not NVFP4 encoding.
  • —64-layer text trunk; 1,391 tensors in the complete package.
  • —192 MLP linear weights use FP8 E4M3 with 128×128 blocks; 305 other text linear weights remain BF16.
  • —vision-mtp-bf16.safetensors is a 1,770,897,648-byte shared component containing the official 333 BF16 vision tensors and 15 BF16 MTP tensors. It is not a standalone main model.
  • —The official MTP component was restored from Qwen revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; it was not trained during this SFT/SimPO run.

<!-- W8A16FINALCAPABILITYV1EN_START -->

W8A16 formal capability results

Protocol: one RTX PRO 6000 96GB, SGLang + official BF16 MTP, xhigh, 32,768-token output limit, and a 1,800-second request timeout. GPQA and LCB used C16; MMLU is the adopted clean C24 result.

SuiteFinal scoreConcurrencyRequest errors / timeouts
GPQA159/198 (80.30%)C160 / 0
MMLU450/500 (90.00%)C240 / 0
LCB78/100 (78.00%)C160 / 0

For LCB, all 100 problems remain in the denominator: 22 length stops and the corresponding 22 empty-code submissions count as failures. The run recorded 0 request errors, 0 HTTP timeouts, 0 code-execution timeouts, and 0 syntax errors. <!-- W8A16FINALCAPABILITYV1EN_END -->

W8A16 MTP short-output concurrency

Measured on one RTX PRO 6000 96GB with SGLang + official MTP, xhigh, and one 256-token wave per cell. All 53/53 requests completed with 0 request errors.

ConcurrencyAggregate tok/sMTP acceptanceAccepted draft / verificationMean TTFT
C18772.9%2.190.063 s
C4 (highest acceptance)32675.8%2.270.143 s
C856275.7%2.270.170 s
C1695574.7%2.240.283 s
C24 (max throughput)1,05673.2%2.200.257 s

This is a single-wave fixed-length sweep, not per-user speed or sustained 32K throughput. Spark SGLang/vLLM bare+MTP text/image/video smoke and PRO SGLang bare+MTP capability smoke also passed; these are short runtime compatibility checks, not formal general-quality scores. The adopted final GPQA/MMLU/LCB scores are listed above; no partial score is reported here.

Tested SGLang path
bash
SGLANG_FORCE_FP8_MARLIN=1 python -m sglang.launch_server \
  --model-path ./W8A16 \
  --quantization modelopt_mixed \
  --dtype bfloat16 --kv-cache-dtype bfloat16 \
  --enable-linear-replayssm-spec \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4

The published metadata view above is the SGLang-tested path. vLLM bare+MTP smoke was validated only through a separate generic-FP8 metadata view of the same unchanged weights, with forced FP8 Marlin; that auxiliary view is not part of this download, so SGLang is the recommended published path. <!-- W8A16RELEASEV1ENEND -->

MTP concurrency, speed and acceptance

This is the existing 256-token fixed-length short-output sweep: one wave per cell, 53 requests per variant. Aggregate throughput is neither per-user speed nor sustained 32K reasoning performance. The C16 cell is a completed short-output speed measurement, not a C16 full capability score.

Fast

ConcurrencyAggregate tok/sMTP acceptanceAccepted draft / verificationCompletion / verification
C18275.64%2.2693.282
C431876.56%2.2973.303
C868874.49%2.2353.225
C161,19674.25%2.2273.223
C241,68473.17%2.1953.197

Mixed precision

ConcurrencyAggregate tok/sMTP acceptanceAccepted draft / verificationCompletion / verification
C110272.92%2.1883.200
C437473.08%2.1933.180
C866572.21%2.1663.156
C161,17672.71%2.1813.188
C241,59573.42%2.2033.202

Both measured short-output grids reach their highest aggregate throughput at C24. C24 long-reasoning runs recorded timeouts, so C24 is not recommended as a validated long-output optimum. Acceptance is accepted / proposed draft tokens; it differs from accepted draft length and final completion length per verification. Dividing completion length by 4 does not give acceptance.

Choose and download

DirectoryQuantizationComplete package size
W4A4/W4A4 NVFP4, retained BF16 head/control components20.62 GB
W4A4+W8A8/W4A4 NVFP4 + W8A8 FP8, retained BF16 head/control components25.80 GB
W4A16/W4A16 ModelOpt NVFP4; BF16 activation/KV and official BF16 vision/MTP20.62 GB

Choose mixed precision when prioritizing the capability scores in this evaluation. Fast is smaller and has higher throughput in this C24 short-output sweep. Package size is not VRAM usage. All four are safetensors packages, not GGUF; vision and native MTP remain BF16 rather than NVFP4. Each variant additionally bundles an optional true static FP8 DFlash2 draft under DFlash2-FP8/; DFlash2 and native MTP are alternative speculative decoders.

Download all 15 files in the chosen directory: model-nvfp4-fast.safetensors (W4A4/) or model-nvfp4-mixed.safetensors (W4A4+W8A8/), vision-mtp-bf16.safetensors, configs/index/tokenizer/processors, manifest.json, and SHA256SUMS. Do not download only the text shard or mix variant files. Each variant contains 64 text layers, 333 BF16 vision tensors, and 15 BF16 MTP tensors. The official MTP was not trained during this SFT/SimPO run.

<!-- LYNNAGENTV0867ENSTART --> <!-- AWQW4A16V1ENSTART -->

AWQ-W4A16 | native MTP with vLLM

[image]

Complete directory: AWQ-W4A16/; the main weight is Qwen3.8-27B-EfficientThink-SimPO-AWQ-W4A16.safetensors. Download the entire directory; do not mix it with W4A4, W4A4+W8A8, the earlier W4A16 package, or W8A16.

Format identity: this is a compressed-tensors, pack-quantized AWQ W4A16 build. 367 target weights use asymmetric group-128 INT4, 33 target weights use symmetric group-128 INT8, and critical, vision, MTP, and other retained tensors remain BF16; activations and KV cache are BF16. It is not NVFP4 encoding, GPTQ, or imatrix. hf_quant_config.json is retained as upstream ModelOpt provenance; runtime loading follows config.json, where quant_method=compressed-tensors is authoritative.

vision-mtp-bf16.safetensors combines 333 official BF16 vision tensors and 15 official BF16 MTP tensors. It is the vision/MTP component for this model, not a DFlash2 draft.

Formal capability and reasoning results

Protocol: one RTX PRO 6000 96GB, vLLM + native MTP, C24, xhigh, a 32,768-token output cap, and a 1,800-second request timeout. Every anomalous sample remains in the denominator; anomaly categories may overlap.

SuiteFinal scoreMean reasoningP50 / P90>8K / >16K32K trunc.Request errors / HTTP timeouts
GPQA159/198 (80.30%)11,1426,283 / 32,76887/198 (43.94%) / 56/198 (28.28%)28/198 (14.14%)0/198 / 0/198
MMLU453/500 (90.60%)935218 / 1,674.412/500 (2.40%) / 4/500 (0.80%)0/5000/500 / 0/500
LCB74/100 (74.00%)13,9228,207.5 / 32,76850/100 (50.00%) / 42/100 (42.00%)23/100 (23.00%)0/100 / 0/100

GPQA has 28/198 (14.14%) empty-final, no-submission, and unparseable cases. MMLU has 0/500 empty, no-submission, and unparseable cases. For LCB, 22/100 (22.00%) no-code/no-submission cases, 23/100 (23.00%) unparseable outputs, and 1/100 (1.00%) syntax error remain failures; code-execution timeouts were 0/100. This vLLM C24 run is not a controlled quantization-loss comparison against the earlier SGLang W4A16, NVFP4, or other decoder results.

vLLM MTP short-output concurrency

The fixed workload used 1,024 input + 256 output tokens, three trials per cell, and num_speculative_tokens=3; all request-error counts were 0. TPS is rounded to whole tokens/s.

ConcurrencyAggregate tok/sMTP acceptance
C13956.03%
C412851.82%
C822652.19%
C1636649.05%
C24 (highest measured throughput)48053.46%

Verified launch path

bash
python -m vllm.entrypoints.openai.api_server \
  --model ./AWQ-W4A16 \
  --served-model-name qwen38-27b-awq-w4a16 \
  --host 127.0.0.1 --port 19540 \
  --dtype bfloat16 \
  --quantization compressed-tensors \
  --kv-cache-dtype auto \
  --gpu-memory-utilization 0.95 \
  --max-model-len 65536 \
  --max-num-seqs 24 \
  --max-num-batched-tokens 2048 \
  --reasoning-parser qwen3 \
  --attention-backend TRITON_ATTN \
  --limit-mm-per-prompt '{"image":2,"video":1}' \
  --skip-mm-profiling \
  --no-enable-prefix-caching \
  --enforce-eager \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

vLLM 0.28.0 with compressed-tensors 0.17.0 and Transformers 5.15.1 passed bare/MTP text, Chinese, code, explanation, image, and real-MP4 smoke, including native MTP accepted/proposed counters. SGLang 0.5.19 with compressed-tensors 0.18.0, Transformers 5.12.1, FlashInfer 0.6.18, and Decord 0.6.0 also passed 6/6 in bare and MTP modes, but only inside an isolated overlay with a local CUDA libcudart link repair and an explicit Decord video-backend patch; this must not be read as stock-pip, drop-in compatibility. That SGLang run was a short smoke and provides no SGLang long-output quality or TPS claim. Reproducibility files are under AWQ-W4A16/runtime/. <!-- AWQW4A16V1ENEND -->

<!-- AWQW4A16V1ZHSTART -->

AWQ-W4A16|vLLM 原生 MTP

[image]

完整目录:AWQ-W4A16/;主权重文件为 Qwen3.8-27B-EfficientThink-SimPO-AWQ-W4A16.safetensors。请下载整个目录,勿与 W4A4、W4A4+W8A8、旧 W4A16 或 W8A16 文件混用。

格式身份:这是 compressed-tensors 的 pack-quantized AWQ W4A16:367 个目标权重采用 group-128 非对称 INT4,33 个目标权重采用 group-128 对称 INT8,其余关键、视觉及 MTP 张量保留 BF16;激活与 KV cache 为 BF16。它不是 NVFP4 编码、GPTQ 或 imatrix。hf_quant_config.json 仅保留上游 ModelOpt 来源记录,运行时以 config.json 的 quant_method=compressed-tensors 为准。

vision-mtp-bf16.safetensors 同时包含 333 个官方 BF16 视觉张量和 15 个官方 BF16 MTP 张量;它是主模型的视觉/MTP 组件,不是 DFlash2 draft。

正式能力与思考量

口径:单张 RTX PRO 6000 96GB、vLLM + 原生 MTP、C24、xhigh、32,768 token 输出上限、1,800 秒请求超时。全部异常样本保留在分母;异常项可重叠。

项目最终得分平均思考P50 / P90>8K / >16K32K 截断请求错误 / HTTP 超时
GPQA159/198(80.30%)11,1426,283 / 32,76887/198(43.94%)/ 56/198(28.28%)28/198(14.14%)0/198 / 0/198
MMLU453/500(90.60%)935218 / 1,674.412/500(2.40%)/ 4/500(0.80%)0/5000/500 / 0/500
LCB74/100(74.00%)13,9228,207.5 / 32,76850/100(50.00%)/ 42/100(42.00%)23/100(23.00%)0/100 / 0/100

GPQA 的空 final、无提交和不可解析均为 28/198(14.14%)。MMLU 的空答、无提交和不可解析均为 0/500。LCB 的无代码/无提交为 22/100(22.00%)、不可解析 23/100(23.00%)、语法错误 1/100(1.00%)、代码执行超时 0/100;均按失败计入。这里与旧 SGLang W4A16、NVFP4 或其他解码器结果不是同条件量化损失对照。

vLLM MTP 短输出并发

固定口径为 1,024 输入 + 256 输出,每档 3 次,num_speculative_tokens=3,所有请求错误为 0。TPS 按模型卡统一取整。

并发聚合 tok/sMTP 接受率
C13956.03%
C412851.82%
C822652.19%
C1636649.05%
C24(已测最高吞吐)48053.46%

已验证启动方式

bash
python -m vllm.entrypoints.openai.api_server \
  --model ./AWQ-W4A16 \
  --served-model-name qwen38-27b-awq-w4a16 \
  --host 127.0.0.1 --port 19540 \
  --dtype bfloat16 \
  --quantization compressed-tensors \
  --kv-cache-dtype auto \
  --gpu-memory-utilization 0.95 \
  --max-model-len 65536 \
  --max-num-seqs 24 \
  --max-num-batched-tokens 2048 \
  --reasoning-parser qwen3 \
  --attention-backend TRITON_ATTN \
  --limit-mm-per-prompt '{"image":2,"video":1}' \
  --skip-mm-profiling \
  --no-enable-prefix-caching \
  --enforce-eager \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

vLLM 0.28.0 + compressed-tensors 0.17.0 + Transformers 5.15.1 已完成 bare/MTP 的文本、中文、代码、解释、图片和真实 MP4 smoke,并记录原生 MTP accepted/proposed 计数。SGLang 0.5.19 + compressed-tensors 0.18.0 + Transformers 5.12.1 + FlashInfer 0.6.18 + Decord 0.6.0 的 bare/MTP 同样各完成 6/6,但依赖隔离 overlay、局部 CUDA libcudart 链接修复与显式 Decord 视频后端补丁,不能理解为原生 pip 环境即装即用。SGLang 该轮只是短 smoke;没有 SGLang 长输出质量或 TPS 结论。可复现文件见 AWQ-W4A16/runtime/。 <!-- AWQW4A16V1ZHEND -->

Lynn Agent v0.86.7

Lynn Agent v0.86.7 now uses this release's Q2-LynnStyle / Q3-LynnStyle + DFlash2 packages. The pairing passed runtime validation on DGX Spark; macOS notarization, CI in both repositories, synchronization across all three release repositories, and public-network SHA verification also passed. Release-cache cleanup reclaimed approximately 4.01 GB, leaving approximately 22.34 GB free on the Tencent mirror disk; the current build, previous build, model files, and running service were not changed.

InstallerChina mirrorGitHub fallback
Mac Apple SiliconDownloadDownload
Mac IntelDownloadDownload
WindowsDownloadDownload

Release records: primary GitHub repository · legacy GitHub repository · Gitee · CLI package <!-- LYNNAGENTV0867ENEND -->

Known boundaries

C24 observationFastMixed precision
GPQA request timeouts1211
GPQA output-limit returns / empty final among returned requests14 / 1413 / 13
MMLU request errors / empty finals0 / 00 / 0
LCB request errors00
LCB output-limit returns / empty code25 / 2625 / 25

These counters overlap and must not be added as disjoint failures. Scores follow the frozen scorer's final decisions; an empty visible final is a separate observation. We do not claim zero strict loops or losslessness in every mode. A fast-variant non-thinking code probe answered incorrectly in both bare and MTP modes; its XH counterpart passed. The scored SGLang C24 evidence still covers the MTP path only; the separate vLLM 0.28.0 bare/MTP four-mode smoke is documented below.

Training method

Official Qwen3.8-27B → capability-preserving SFT → terminal-behavior SimPO → per-tensor FP32 delta merge → BF16 → three NVFP4 exports.

  • —SFT: 1,905 samples, 1 epoch, 239 optimizer steps, effective batch 8; LoRA r16 / alpha32 / dropout 0.05; LR 5e-6 and 12 warmup steps.
  • —SimPO: 110 preference pairs, 73 unique prompts, 5 optimizer steps; beta 1.0, gamma 0.2, peak LR 5e-7; LoRA r16 / alpha32 / dropout 0; two-GPU FSDP full sharding.
  • —Training hardware: 2× RTX PRO 6000 Blackwell Server Edition.

SGLang runtime notes

SGLang native MTP retains the formal C24 quality and throughput evidence; vLLM 0.28.0 now has a separate four-mode startup/load/generation smoke. The frozen Spark launch records use the settings below; do not interchange the quantization backends:

SettingFastMixed precision
--quantizationmodelopt_fp4modelopt_mixed
--speculative-algorithmEAGLEEAGLE
steps / top-k / draft tokens3 / 1 / 43 / 1 / 4
KV / contextBF16 / 65,536BF16 / 65,536

That frozen environment uses SGLang source revision 17313cf4b25d with runtime adaptations; this is not a clean-upstream-install compatibility guarantee. The formal scores and short sweeps above were measured on PRO 6000, not Spark. Do not reuse GGUF DFlash2 arguments or private container-image names.

This page retains the C24 capability report.

<!-- VLLM028RUNTIMEV1ENSTART -->

vLLM 0.28.0 · tested on DGX Spark

Fast and mixed precision both completed real bare and native-MTP load, health, and text/image/video generation checks. The environment was DGX Spark GB10 (SM121) with the official Linux/ARM64 vLLM 0.28.0 image. Fast used modelopt_fp4; mixed precision used modelopt_mixed. Both selected FlashInferCutlassNvFp4LinearKernel for NVFP4, while mixed FP8 layers selected FlashInferFP8ScaledMMLinearKernel; neither fell back to Marlin.

VariantDecodeLoad memoryLoad time6-case content checkImage / videoMTP accepted / drafted tokens
Fastbare18.77 GiB145.10 s5/6pass / pass—
FastMTP19.56 GiB198.43 s5/6pass / pass101/168 (60.1%)
Mixed precisionbare23.52 GiB143.88 s5/6pass / pass—
Mixed precisionMTP24.31 GiB211.90 s6/6pass / pass112/168 (66.7%)

All six requests in every mode returned HTTP 200 with non-empty output and no repetitive-punctuation collapse. One short code check returned 55 instead of the expected 30 in fast bare/MTP and mixed bare, so those modes are reported as 5/6; mixed-precision MTP was 6/6. This was a short serial non-thinking smoke, not a vLLM TPS benchmark, general quality proof, 64K long-context test, or concurrency stress test. The 65,536 context and max-num-seqs=4 values were startup settings.

The recommended vLLM starting point is mixed precision + native MTP:

bash
VARIANT=quality MODE=mtp PORT=19120 bash runtime/START_VLLM028_NVFP4.sh

The launcher binds only to 127.0.0.1, pins method=mtp and num_speculative_tokens=3, and refuses to pull an image or overwrite an existing container. Preload the image first; the script verifies the pinned official immutable manifest and local image ID. Model, quantization, and MTP arguments match the smoke above. The loopback host-network mapping is a deployment adapter for host access, not a new performance run.

Official references: vLLM ModelOpt quantization, vLLM MTP, and Docker host networking. <!-- VLLM028RUNTIMEV1ENEND --> <!-- NVFP4FILENAMEV4ENSTART --> <!-- NVFP4FILENAMEV4ENEND --> <!-- NVFP4CARDV2ENEND -->


<!-- NVFP4CARDV2ZHSTART -->

Qwen3.8-27B EfficientThink — NVFP4 + BF16 MTP

[image]

BF16 / FP8 主仓 · GGUF 版本

<!-- DFLASHSTATICFP8V1ZH_START -->

名副其实的静态 FP8 DFlash2 draft

仓内 DFlash2 draft 现已替换为预量化静态 FP8 compressed-tensors checkpoint,不再是放在 FP8 目录名下的 BF16 文件。model.safetensors 为 2,407,027,720 bytes(SHA256 1f3636a32d866f8ebc7f422d63f9247126ebb6d2566d3e0da327d81dd8fa25d1);tensor 审计为 20 个 FP8 E4M3 权重、20 个 FP32 scale,以及 61 个保留 BF16 tensor。

加载时必须显式加入:

text
--speculative-draft-model-quantization compressed-tensors

同条件 DGX Spark 对照使用同一 W4A4 主模型、15 条固定输入、XH、256 输出 token 与 8 draft tokens;四个 cell 均为 15/15 请求成功。

Draft并发聚合 tok/sDFlash 接受率平均接受长度 / 8请求错误
BF16 对照C129.7538.77%3.720
静态 FP8C130.2635.86%3.510
BF16 对照C468.5032.71%3.290
静态 FP8C475.9433.55%3.350

C1 接受度为同轮服务日志快照均值,C4 接受度为结束后 SGLang metrics 精确值。本次同口径 C4 中,静态 FP8 draft 的聚合吞吐比 BF16 高约 10.9%。该短定长测试只验证服务行为,不替代本卡其他位置的正式能力分数。 <!-- DFLASHSTATICFP8V1ZH_END -->

科研免责声明: 本实验版本仅用于研究拒答倾向消融的技术可行性及行为影响,不构成全面的安全结论、对无限制使用的认可或专业建议。用户应依法、恰当地使用,并独立核验模型输出。

图中并列展示本模型的 W4A4、W4A4+W8A8、W4A16 与 W8A16。Fast/Mixed 为 C24,W4A16 为 C16;W8A16 的 GPQA/LCB 为 C16、MMLU 为 C24。GPQA 的 12 / 11 个超时均保留在 Fast/Mixed 的 198 题失败分母中,未丢弃异常题。

<!-- INT8W8A8QATV2ZH_START -->

真 QAT INT8 W8A8|动态 INT8 激活

[image]

下载目录:`INT8-W8A8-QAT/`。 同平台对应主仓/独立仓。

文件与精度

组件路径精度 / 角色大小
主模型INT8-W8A8-QAT/model-00001-of-00008.safetensors … model-00008-of-00008.safetensors真 QAT INT8 W8A8;动态 INT8 激活29.48 GB
视觉 + 原生 MTPINT8-W8A8-QAT/vision-mtp-bf16.safetensors333 个 BF16 视觉张量 + 15 个 BF16 MTP 张量,索引真实引用 348 项1.77 GB
完整清单INT8-W8A8-QAT/manifest.json 与 INT8-W8A8-QAT/SHA256SUMS当前目录 29 个文件的角色、bytes 与 SHA256—
结构化评测`INT8-W8A8-QAT/evaluation/formal-quality-and-performance.json`正式成绩、思考量与完整研究记录—

训练与导出方法

  • —64 层 Qwen3.8-27B 多模态架构,发布包索引共 1,599 个张量。
  • —经过 3,200 个 QAT optimizer steps;400 个语言线性层均记录到非零梯度,并以 INT8 权重发布。
  • —激活采用动态 INT8;247 条训练样本进入已接受训练集。
  • —BF16 scale 无损导出为 F32;视觉塔与原生 MTP 保持 BF16。

正式能力与思考量

协议:单张 RTX PRO 6000 Blackwell 96GB、vLLM 0.28.0 + 原生 MTP3、C20、BF16 KV、reasoning_effort=xhigh、32,768 输出上限、1,800 秒请求超时。正式长输出采用 C20,因为 C24 无法为整套长输出保留足够 KV 容量。

项目得分平均思考P50 / P90>8K / >16K32K 截断空 final / 不可解析
GPQA162/198(81.82%)9,5304,699.5 / 32,76773 / 412323 / 23
MMLU451/500(90.20%)837203 / 1,906.19 / 300 / 0
LCB73/100(73.00%)13,9218,347 / 32,76850 / 412323 / 23

请求 / HTTP / capture / grader 错误均为 0;LCB 代码超时与语法错误均为 0,另有 1 个 runtime error。LCB 采用 IPC-v4 对原始 100 份回答统一复判,没有发出新模型请求。

24 个随包运行路径的短输出性能单元

协议:1,024 输入 / 256 输出、预热后 3 轮。下表只代表短定长服务吞吐,不代表长思考速度;24 个裸跑/MTP3 单元均为 0 请求错误。

  • —本档最高实测吞吐:vLLM MTP3 C24,662 tok/s,接受率 56.28%,约 27.6 tok/s/请求。
  • —SGLang MTP3 C24:654 tok/s,接受率 55.17%。
  • —发布包只附带并推荐原生 MTP3;完整历史研究数据保留在结构化评测文件中。
框架 / 模式并发聚合 tok/s每请求 tok/s接受率TTFT P50延迟 P50峰值显存错误
vLLM 裸跑C13231.9—0.16s8.03s85.8 GiB0
vLLM 裸跑C411328.2—0.56s9.04s86.1 GiB0
vLLM 裸跑C821326.6—1.01s9.55s86.1 GiB0
vLLM 裸跑C1636322.7—1.59s11.15s86.1 GiB0
vLLM 裸跑C2042321.2—1.87s11.95s86.1 GiB0
vLLM 裸跑C2447619.8—2.16s12.71s86.1 GiB0
vLLM MTP3C15655.947.48%0.18s4.58s85.8 GiB0
vLLM MTP3C420250.454.87%0.56s4.62s86.0 GiB0
vLLM MTP3C835644.557.08%1.09s5.27s86.0 GiB0
vLLM MTP3C1654434.057.31%1.70s6.88s86.0 GiB0
vLLM MTP3C2061931.056.30%2.00s7.77s86.0 GiB0
vLLM MTP3C2466227.656.28%2.31s8.63s86.0 GiB0
SGLang 裸跑C14545.2—0.14s5.67s87.9 GiB0
SGLang 裸跑C415839.4—0.43s6.49s88.1 GiB0
SGLang 裸跑C828735.9—0.69s7.13s88.1 GiB0
SGLang 裸跑C1646929.3—1.22s8.73s88.1 GiB0
SGLang 裸跑C2053726.8—1.48s9.53s88.1 GiB0
SGLang 裸跑C2459724.9—1.74s10.29s88.1 GiB0
SGLang MTP3C18686.167.86%0.15s2.97s86.5 GiB0
SGLang MTP3C423458.554.03%0.43s3.83s86.7 GiB0
SGLang MTP3C837947.451.95%0.72s4.96s86.7 GiB0
SGLang MTP3C1657836.155.06%1.25s6.56s86.7 GiB0
SGLang MTP3C2061230.654.42%1.52s8.06s86.7 GiB0
SGLang MTP3C2465427.355.17%1.79s8.90s86.7 GiB0

已验证启动方式

bash
cd INT8-W8A8-QAT
bash scripts/serve-vllm-mtp3.sh
脚本用途
scripts/serve-vllm-bare.shvLLM 裸跑
scripts/serve-vllm-mtp3.shvLLM 原生 MTP3;推荐吞吐路径
scripts/serve-sglang-bare.shSGLang 裸跑
scripts/serve-sglang-mtp3.shSGLang 原生 MTP3

四条随包路径均在同哈希模型上通过文本、图片与真实视频 smoke。Blackwell SM120/121 的 SGLang compressed-tensors INT8 路径使用随包 runtime/sglang-sm120-int8-compat/ 兼容层。

<!-- INT8W8A8QATV2ZH_END -->

三项正式成绩

版本GPQA 198MMLU 500LCB 100
W4A4(NVFP4 极速版)158/198(79.80%)447/500(89.40%)74/100(74.00%)
W4A4+W8A8(NVFP4 混合精度版)168/198(84.85%)458/500(91.60%)75/100(75.00%)
W4A16161/198(81.31%)457/500(91.40%)74/100(74.00%)
W8A16159/198(80.30%)450/500(90.00%)78/100(78.00%)

Fast/Mixed 的三项成绩均使用 C24 运行协议。GPQA 的 12 / 11 个超时保留在完整分母中;未剔除或重算异常题。不同运行格式及解码器的结果不能直接当作等条件量化损失。

<!-- NVFP4BENEFITSV4ZHSTART -->

混合精度版相对极速版:得分与思考量

指标W4A4 极速版W4A4 + W8A8 混合精度版变化
GPQA 得分158/198 (79.80%)168/198 (84.85%)+10题 / +5.05pp
GPQA 平均思考量9,1238,400-7.9%
GPQA P50 / P904,928.5 / 24,893.04,278.0 / 24,402.6—
MMLU 得分447/500 (89.40%)458/500 (91.60%)+11题 / +2.20pp
MMLU 平均思考量963848-11.9%
MMLU P50 / P90225.0 / 2,285.5216.5 / 1,668.9—
LCB 得分74/100 (74.00%)75/100 (75.00%)+1题 / +1.00pp
LCB 平均思考量14,60213,647-6.5%
LCB P50 / P909,626.5 / 32,769.07,441.5 / 32,769.0—

同一轮 C24 协议下,混合精度版 GPQA 多答对 10 题、MMLU 多答对 11 题、LCB 多答对 1 题;三项平均思考 token 均减少。这里比较的是量化版本,不是原版与后训练版。MMLU/LCB 均值分母为 500/100,百分比变化从未四舍五入的均值计算。 <!-- NVFP4BENEFITSV4ZHEND -->

<!-- W4A16FINALV1ZHSTART -->

W4A16 正式 C16 全量结果

协议:单张 RTX PRO 6000 96GB、SGLang + 官方 BF16 MTP、C16、xhigh、32,768 输出上限、1,800 秒请求超时。此处与 W4A4 两版的 C24 不是严格等并发对照。

项目最终得分平均思考P50 / P90>8K / >16K32K 截断其他异常
GPQA161/198(81.31%)11,4046,795.5 / 32,76792 / 5727final 通道空 29;不可解析 27;请求错误 0
MMLU457/500(91.40%)802225 / 1,5549 / 10请求错误 0;空 final 0
LCB74/100(74.00%)14,5118,513.5 / 32,76951 / 4226请求错误/超时 0;空代码 26

<!-- W4A16GPQABOUNDARYV1ZH_START -->

GPQA 计分:全部 198 题的最终成绩为 161/198(81.31%),请求错误 0、超时 0。解析仅接受非空 final 通道,或 reasoning 最后一个非空行中的完整明确 Final Answer: A/B/C/D。

<!-- W4A16GPQABOUNDARYV1ZH_END -->

LCB 难度:Easy 23/23(100%)、Medium 29/31(93.55%)、Hard 22/46(47.83%)。26 个长度结束与 26 个空代码均按失败保留在 100 题分母。

三套 NVFP4 的 MMLU 均采用同一 final-content-strict-single-letter-v2 离线重算:极速版 447/500(89.40%)、混合精度版 458/500(91.60%)、W4A16 457/500(91.40%);生成内容未改变。旧的 430/449/439 分数不再使用。

W4A16 MTP 短输出并发

固定短输出网格仅测试 C1/C4/C8/C16/C24;C2 与 C32 未测。聚合吞吐按模型卡规则取整。

并发聚合 tok/sMTP 接受率接受草稿 / 验证轮实际提交 / 验证轮
C112772.84%2.1853.160
C445270.20%2.1063.103
C876471.04%2.1313.127
C16(平衡推荐)1,21674.28%2.2293.228
C24(最高吞吐)1,30470.60%2.1183.114

五档均为 0 请求错误、0 超时、0 空输出;本短扫未进行标点坍塌人工终审,因此不作对应零值声明。C16 在保持 1,216 tok/s 时接受率最高,作为平衡档;C24 是已测最大聚合吞吐。 <!-- W4A16FINALV1ZHEND -->

W4A16 的已测 SGLang 参数为 --quantization modelopt_mixed、EAGLE、steps=3、top-k=1、draft tokens=4、BF16 dtype/KV、65,536 context;vLLM 四路 smoke 仍只覆盖极速版与混合精度版。

<!-- W8A16RELEASEV1ZHSTART -->

W8A16 质量优先 FP8 包

完整 W8A16 包已发布在 W8A16/:含清单共 24 个文件 / 38,477,562,434 bytes。请下载整个目录,不要与 W4A4/、W4A4+W8A8/ 或 W4A16/ 混用。

格式说明: W8A16 使用分块 FP8 E4M3 权重 + BF16 激活与 KV cache。它为了统一分发放在本仓系列中,但编码格式不是 NVFP4。
  • —64 层文字主干;完整包共 1,391 个张量。
  • —192 个 MLP 线性权重使用 128×128 分块 FP8 E4M3;其余 305 个文字线性权重保留 BF16。
  • —vision-mtp-bf16.safetensors 为 1,770,897,648-byte 共用组件,包含官方 333 个 BF16 视觉张量与 15 个 BF16 MTP 张量;它不是可独立运行的主模型。
  • —官方 MTP 来自 Qwen revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0,未参与本轮 SFT/SimPO 训练。

<!-- W8A16FINALCAPABILITYV1ZH_START -->

W8A16 正式能力成绩

口径:单张 RTX PRO 6000 96GB、SGLang + 官方 BF16 MTP、xhigh、32,768 token 输出上限、1,800 秒请求超时。GPQA 与 LCB 使用 C16;MMLU 采用已冻结的 clean C24 结果。

项目最终得分并发请求错误 / 超时
GPQA159/198(80.30%)C160 / 0
MMLU450/500(90.00%)C240 / 0
LCB78/100(78.00%)C160 / 0

LCB 保留全部 100 题作为分母:22 次长度结束及对应的 22 个空代码均按失败计入;请求错误、HTTP 超时、代码执行超时和语法错误均为 0。 <!-- W8A16FINALCAPABILITYV1ZH_END -->

W8A16 MTP 短输出并发

实测环境:单张 RTX PRO 6000 96GB、SGLang + 官方 MTP、xhigh、每档单波 256-token 定长输出。53/53 个请求全部完成,请求错误 0。

并发聚合 tok/sMTP 接受率接受草稿 / 验证轮平均 TTFT
C18772.9%2.190.063 秒
C4(接受率最高)32675.8%2.270.143 秒
C856275.7%2.270.170 秒
C1695574.7%2.240.283 秒
C24(最高吞吐)1,05673.2%2.200.257 秒

这是单波定长短测,不是单用户速度,也不是 32K 持续吞吐。Spark 上的 SGLang/vLLM bare+MTP 文本/图片/视频 smoke,以及 PRO 上的 SGLang bare+MTP 能力 smoke 也已通过;它们只证明短请求运行兼容性,不是正式通用质量分数。已采用的 GPQA/MMLU/LCB 最终成绩见上方;本段不披露 partial 分数。

已测 SGLang 路径
bash
SGLANG_FORCE_FP8_MARLIN=1 python -m sglang.launch_server \
  --model-path ./W8A16 \
  --quantization modelopt_mixed \
  --dtype bfloat16 --kv-cache-dtype bfloat16 \
  --enable-linear-replayssm-spec \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4

以上已发布元数据是 SGLang 实测路径。vLLM bare+MTP smoke 仅通过同一权重的独立通用 FP8 元数据视图并强制 FP8 Marlin 完成;该辅助视图未随本目录发布,因此公开推荐路径仍为 SGLang。 <!-- W8A16RELEASEV1ZHEND -->

MTP 并发速度与接受率

以下是既有256-token 定长短输出测试,每档单批,两版各 53 个请求。聚合吞吐不是单用户速度,也不是 32K 长推理持续吞吐。这里的 C16 是已完成的短输出并发测速,不代表 C16 全量能力评分。

极速版

并发聚合 tok/sMTP 接受率接受草稿 / 验证轮实际输出 / 验证轮
C18275.64%2.2693.282
C431876.56%2.2973.303
C868874.49%2.2353.225
C161,19674.25%2.2273.223
C241,68473.17%2.1953.197

混合精度版

并发聚合 tok/sMTP 接受率接受草稿 / 验证轮实际输出 / 验证轮
C110272.92%2.1883.200
C437473.08%2.1933.180
C866572.21%2.1663.156
C161,17672.71%2.1813.188
C241,59573.42%2.2033.202

两版已测短输出网格均在 C24 取得最高聚合吞吐,但 C24 长推理出现过超时,因此暂不推荐它作为长输出最佳并发。接受率按实际接受草稿数 / 提议草稿数计算;它与每轮接受长度、每轮最终输出不同,不能用输出长度除以 4 替代。

选择与下载

目录量化方案完整包大小
W4A4/W4A4 NVFP4,保留 BF16 头部与控制部件20.62 GB
W4A4+W8A8/W4A4 NVFP4 + W8A8 FP8,保留 BF16 头部与控制部件25.80 GB
W4A16/W4A16 ModelOpt NVFP4;BF16 activation/KV,官方 BF16 视觉与 MTP20.62 GB

混合精度版更适合优先考虑本轮能力成绩的使用者;极速版文件较小,在本次 C24 短输出下吞吐更高。文件大小不是显存需求。四版都是 safetensors 包,不是 GGUF;视觉与原生 MTP 均为 BF16,而非 NVFP4。每个档位另在 DFlash2-FP8/ 提供可选的真静态 FP8 DFlash2 draft;DFlash2 与原生 MTP 是两条可选的投机解码路径。

下载所选目录的完整 15 个文件:model-nvfp4-fast.safetensors(W4A4/)或 model-nvfp4-mixed.safetensors(W4A4+W8A8/)、vision-mtp-bf16.safetensors、配置/索引/tokenizer/processor、manifest.json 和 SHA256SUMS。不要只下载文字分片,也不要把三个目录的文件混放。每版含 64 层文字主干、333 个 BF16 视觉张量及 15 个 BF16 MTP 张量;官方 MTP 未参与本次 SFT/SimPO 训练。

<!-- LYNNAGENTV0867ZHSTART -->

Lynn Agent v0.86.7

Lynn Agent v0.86.7 已采用本系列 Q2-LynnStyle / Q3-LynnStyle + DFlash2。组合已在 DGX Spark 实测通过;Mac 公证、两仓 CI、三仓同步与公网 SHA 校验均通过。发布缓存清理后实际释放约 4.01 GB,腾讯镜像盘剩余约 22.34 GB;当前版、上一版、模型文件和服务均未改动。

安装包国内镜像GitHub 备用
Mac Apple Silicon下载下载
Mac Intel下载下载
Windows下载下载

发布记录:GitHub 主仓 · GitHub 旧仓 · Gitee · CLI 包 <!-- LYNNAGENTV0867ZHEND -->

已知边界

C24 观测极速版混合精度版
GPQA 请求超时1211
GPQA 达到输出上限 / 已返回请求空 final14 / 1413 / 13
MMLU 请求错误 / 空 final0 / 00 / 0
LCB 请求错误00
LCB 达到输出上限 / 空代码25 / 2625 / 25

同一题可同时进入多个计数,不能相加为互斥失败数。得分采用冻结判分器最终结果,空可见 final 是单独的观测项。当前不宣称严格死循环为零或全模式无损。极速版不思考代码探针曾在 bare 和 MTP 模式均答错,XH 对应探针答对。SGLang C24 计分仍只对应 MTP 路径;另行完成的 vLLM 0.28.0 bare/MTP 四路 smoke 见下方。

训练方法

官方 Qwen3.8-27B → 能力保持 SFT → 终止行为 SimPO → 逐张量 FP32 增量合并 → BF16 → 三版 NVFP4 导出。

  • —SFT:1,905 条样本,1 epoch,239 个优化器步,有效 batch 8;LoRA r16 / alpha32 / dropout 0.05;LR 5e-6、12 步 warmup。
  • —SimPO:110 对偏好数据、73 个唯一 prompt,5 个优化器步;beta 1.0、gamma 0.2、峰值 LR 5e-7;LoRA r16 / alpha32 / dropout 0,双卡 FSDP full sharding。
  • —训练硬件:2× RTX PRO 6000 Blackwell Server Edition。

SGLang 运行说明

SGLang 原生 MTP 保留正式 C24 质量与吞吐数据;vLLM 0.28.0 已另行完成四路启动/加载/生成 smoke。冻结的 Spark 启动记录使用以下参数,不能把两版的量化后端混用:

设置极速版混合精度版
--quantizationmodelopt_fp4modelopt_mixed
--speculative-algorithmEAGLEEAGLE
steps / top-k / draft tokens3 / 1 / 43 / 1 / 4
KV / contextBF16 / 65,536BF16 / 65,536

该冻结环境使用 SGLang 源码版本 17313cf4b25d,并带运行环境适配;它不是干净上游安装的兼容性保证。上方正式得分和短扫则来自 PRO 6000,不能当作 Spark 的测速。不要照搬 GGUF DFlash2 参数或私有镜像名。

本页继续保留 C24 能力报告。

<!-- VLLM028RUNTIMEV1ZHSTART -->

vLLM 0.28.0 · DGX Spark 实测

极速版与混合精度版均完成 bare 与原生 MTP 四路真实加载、健康检查和文本/图片/视频生成。实测环境为 DGX Spark GB10(SM121)、官方 Linux/ARM64 vLLM 0.28.0 镜像;极速版使用 modelopt_fp4,混合精度版使用 modelopt_mixed。三版 NVFP4 均实际选择 FlashInferCutlassNvFp4LinearKernel,混合精度版的 FP8 层使用 FlashInferFP8ScaledMMLinearKernel,未回退到 Marlin。

版本解码加载显存加载时间6 项内容校验图片 / 视频MTP 接受 token / 提议 token
极速版bare18.77 GiB145.10 秒5/6通过 / 通过—
极速版MTP19.56 GiB198.43 秒5/6通过 / 通过101/168(60.1%)
混合精度版bare23.52 GiB143.88 秒5/6通过 / 通过—
混合精度版MTP24.31 GiB211.90 秒6/6通过 / 通过112/168(66.7%)

四路各 6 个请求均为 HTTP 200、非空输出,未出现重复标点坍塌。极速 bare/MTP 与混合 bare 的同一道简短代码校验答为 55,正确值是 30,因此如实记为 5/6;混合精度 MTP 为 6/6。这里是关闭思考的短串行 smoke,不是 vLLM TPS、通用质量、64K 长上下文或并发压力证明;65,536 context 与 max-num-seqs=4 是启动配置值。

推荐的 vLLM 起步组合是混合精度版 + 原生 MTP:

bash
VARIANT=quality MODE=mtp PORT=19120 bash runtime/START_VLLM028_NVFP4.sh

启动器默认只监听 127.0.0.1,固定 method=mtp 与 num_speculative_tokens=3,并拒绝自动拉取镜像或覆盖同名容器。必须预先准备镜像;脚本会核对固定的官方不可变 manifest 与本地 image ID。模型/量化/MTP 参数来自上述实测;便于宿主访问的 loopback host-network 连接是部署适配,不作为新的性能实测。

官方参考:vLLM ModelOpt 量化、vLLM MTP、Docker host network。 <!-- VLLM028RUNTIMEV1ZHEND --> <!-- NVFP4FILENAMEV4ZHSTART --> <!-- NVFP4FILENAMEV4ZHEND --> <!-- NVFP4CARDV2ZHEND -->