SeatownSin/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-W4A16
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-W4A16
Almost every public NVFP4 checkpoint is W4A4. This one is weight-only: 4-bit weights, BF16 activations, no calibration corpus anywhere in its construction — and therefore nothing you would have to reproduce to rebuild it.
20.6 GB. Runs on one GB10 with images and video intact, and keeps the MTP head so speculative decoding works without a separate drafter.
Is this for you
You have an NVIDIA DGX Spark — GB10, SM121, one coherent ~121 GiB memory pool — and you want a 27B multimodal model that leaves room for real context. You feed it screenshots, photos, or video, not just text.
If you are on a discrete GPU, nothing here is wrong, but none of it was tuned for you. The choices below assume unified memory, where "GPU memory" and "host RAM" are the same DRAM and offloading reclaims nothing.
Why weight-only
W4A16 and W4A4 store the same 4-bit weights. The difference is what else the build depends on. W4A4 quantizes activations too, and activation scales have to come from a calibration corpus — the checkpoint inherits whatever that corpus contained, and anyone rebuilding it has to reproduce the corpus to reproduce the model. Weight-only derives every scale from the weights themselves: rebuilding this takes the source checkpoint, quantize_w4a16.py from this repo, and nothing else.
Throughput does not break the tie. Decode on GB10 is bandwidth-bound: both schemes move the same 4-bit bytes, and this build dequantizes FP4 to BF16 rather than touching the FP4 tensor cores at all. We have not measured W4A4 on this box, so no speed comparison is offered here. Published figures for either scheme are usually extrapolated rather than measured.
Does W4A4 degrade image and video, as folklore says? Not tested here. No W4A4 build of this checkpoint was made, so no result can be shown in either direction. Treat the weight-only choice as justified by the calibration argument above, not by any claim that W4A4 fails.
So weight-only is the simpler construction, not a rescue from a broken alternative. The compression is the win; the 4-bit math never was.
Run it
The launcher ships in this repo — there is nothing else to clone.
hf download SeatownSin/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-NVFP4-W4A16 --local-dir ./qwen38-nvfp4
cd qwen38-nvfp4 && chmod +x serve_export.sh
DST=. ./serve_export.shThat gives you an OpenAI-compatible endpoint on :30000, with MTP speculative decoding, 262K context, and --mem-fraction-static 0.75.
If you would rather not run someone else's shell script, the whole thing is one docker run — this is what the script does, minus the readiness polling:
docker run -d --name qwen38-w4a16 --network host --ipc host --gpus all --shm-size 32g \
-v "$PWD:/model:ro" lmsysorg/sglang:qwen38-27b \
python3 -m sglang.launch_server --model-path /model --trust-remote-code \
--mem-fraction-static 0.75 --attention-backend flashinfer \
--chunked-prefill-size 8192 --disable-prefill-cuda-graph \
--kv-cache-dtype fp8_e4m3 --mamba-ssm-dtype bfloat16 \
--mamba-full-memory-ratio 4.21 --mamba-radix-cache-strategy extra_buffer_lazy \
--context-length 262144 --speculative-algorithm EAGLE \
--speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 \
--reasoning-parser qwen3 --tool-call-parser qwen3_coder \
--host 0.0.0.0 --port 30000Do not serve this with MiaAI's `start.sh`. It has noMODEL_PATH:QUANT=nvfp4is hardcoded to a different checkpoint, and only the HF cache is mounted, so it will quietly download and serve someone else's weights while appearing to serve yours. Its tuning notes are worth reading; its launch path is not.
One more thing: do not copy --mem-fraction-static 0.95 from elsewhere. On one coherent pool that starves the operating system as well as the server, and a long video prefill will take the whole box down with a global OOM — measured here, exit 137, dbus killed too.
Measured on the GB10
These are measurements from the machine that built this checkpoint, not projections.
Speculative decoding runs on the checkpoint's own MTP head at depth 3/1/4, the setting serve_export.sh ships. That depth was swept on the sibling GAIN-V1.1 build, not re-swept here. It transfers because the MTP head is the stock Qwen/Qwen3.8-27B head: DavidAU's finetunes leave the mtp.* tensors byte-identical to the base model (verified by blob hash on the GAIN source), so the head is the same weights in both builds. On that sweep 3/1/4 peaked on code and 2/1/3 was marginally better on prose; spec_sweep.sh is in this repo if you want the numbers for this build on your own prompts.
You will see higher figures quoted for GB10. Most are extrapolated from other hardware rather than measured on one — worth checking which you are reading.
What is inside
MLP is about 62% of the parameters, which is why quantizing it and leaving everything sensitive alone gets most of the compression for very little of the risk. Produced with NVIDIA Model Optimizer's stock W4A16_NVFP4_CFG.
One packaging note if you are reading the files: hf_quant_config.json declares MIXED_PRECISION with a per-module map rather than a bare W4A16_NVFP4. That is deliberate. SGLang rejects the latter outright even though it has the matching kernel; the map is the only form that reaches it. ModelOpt's original config is preserved alongside as hf_quant_config.modelopt.json.
How this was verified
Smoke test: 6/6 cases passed (5 vision). Export verification: 1999 tensors, no activation scales, MTP head and vision tower present and unquantized.
Those are structural checks. They cannot tell you whether 4-bit weights cost real capability, because the realistic failure of a good quantization is output that is worse but still coherent, which nothing automatic catches. So:
24 prompts across text, images and video were run against both this checkpoint and BF16 source / SGLang, on the same engine so the weights are the only variable. No fatal flags were raised. Every ground-truth string the reference matched, this build matched too — including all 6 subtitle strings burned into 3 real video clips, read back intact.
If image or video quality matters to you, run that comparison on your inputs. eval_harness.py, make_evalset.py and serve_export.sh are all in this repo.
Lineage
`Qwen/Qwen3.8-27B` → `DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU` → this.
Finetune by DavidAU: the TURBO-Fable Cold Fusion line (multi-stage tuning with Unsloth, first three stages by Nightmedia), released as a Heretic variant — abliterated/uncensored by its author. Read DavidAU's card for what that means for your use; this page only changes the weight format. Base model by Qwen.
Quantized from source commit `bb98243c7d93af192431688cf3763d70c23f46e6` (2026-09-02). The source repo was being updated hourly at the time and its author has said more versions are coming, so pin that commit if you rebuild; main may not be the same weights.
The full quantization and evaluation pipeline is in this repo: quantize_w4a16.py, verify_export.py, smoke_test.py, eval_harness.py, serve_export.sh, spec_sweep.sh. Everything asserted on this page can be re-derived with them.
Apache 2.0, unbroken from the base model.
