CoolFace
Modelpublic

maci0/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
7likes200downloads
Model Card

<div style="font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;max-width:760px;border:1px solid #cec9ba;background:#f2efe6;color:#14130f;padding:24px 26px;border-radius:2px;margin-bottom:26px;"><div style="display:flex;justify-content:space-between;align-items:center;gap:10px 14px;flex-wrap:wrap;"><span style="display:inline-flex;align-items:center;gap:9px;"><svg width="18" height="18" viewBox="0 0 24 24" aria-hidden="true" style="flex-shrink:0;"><path d="M4 4 H16 L20 8 V20 H4 Z" fill="none" stroke="#6f6b60" stroke-width="1.25"/><path d="M16 4 V8 H20" fill="none" stroke="#6f6b60" stroke-width="1.25"/><circle cx="9" cy="13" r="1" fill="#6f6b60"/><circle cx="13" cy="15" r="1" fill="#6f6b60"/></svg><span style="font-family:ui-monospace,SFMono-Regular,Menlo,monospace;font-size:12px;font-weight:700;letter-spacing:0.02em;">RQ-27B-FABLE-U</span></span><span style="font-family:ui-monospace,SFMono-Regular,Menlo,monospace;font-size:10px;font-weight:700;letter-spacing:0.12em;text-transform:uppercase;color:#b5231c;border:1.5px solid #b5231c;border-radius:2px;padding:2px 7px;display:inline-block;transform:rotate(-3deg);">Uncensored</span></div><div style="font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;font-size:26px;font-weight:800;letter-spacing:-0.02em;line-height:1.15;margin:18px 0 8px;color:#14130f;">Fable-Fusion-711 27B · NVFP4</div><div style="font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;font-size:14.5px;line-height:1.5;color:#4a4740;margin-bottom:18px;">27B thinking VL · uncensored (Heretic) · MTP head + vision tower kept bf16.</div><table style="width:100%;border-collapse:collapse;border:0;margin:2px 0 0;font-variant-numeric:tabular-nums;"><tr><th scope="row" style="text-align:left;font-weight:600;padding:9px 8px 9px 0;border:0;border-top:1.5px solid #14130f;border-bottom:1px solid #cec9ba;background:none;color:#4a4740;letter-spacing:0.01em;font-size:13px;font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;">Params</th><td style="padding:9px 0;border:0;border-top:1.5px solid #14130f;border-bottom:1px solid #cec9ba;background:none;text-align:right;font-weight:700;color:#14130f;font-size:14px;font-family:ui-monospace,SFMono-Regular,Menlo,monospace;white-space:nowrap;">27B</td></tr><tr><th scope="row" style="text-align:left;font-weight:600;padding:9px 8px 9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;color:#4a4740;letter-spacing:0.01em;font-size:13px;font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;">Active</th><td style="padding:9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;text-align:right;font-weight:700;color:#14130f;font-size:14px;font-family:ui-monospace,SFMono-Regular,Menlo,monospace;white-space:nowrap;">27B (dense)</td></tr><tr><th scope="row" style="text-align:left;font-weight:600;padding:9px 8px 9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;color:#4a4740;letter-spacing:0.01em;font-size:13px;font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;">Size</th><td style="padding:9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;text-align:right;font-weight:700;color:#14130f;font-size:14px;font-family:ui-monospace,SFMono-Regular,Menlo,monospace;white-space:nowrap;">20 GB</td></tr><tr><th scope="row" style="text-align:left;font-weight:600;padding:9px 8px 9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;color:#4a4740;letter-spacing:0.01em;font-size:13px;font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;">Perplexity</th><td style="padding:9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;text-align:right;font-weight:700;color:#14130f;font-size:14px;font-family:ui-monospace,SFMono-Regular,Menlo,monospace;white-space:nowrap;">n/a</td></tr><tr><th scope="row" style="text-align:left;font-weight:600;padding:9px 8px 9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;color:#4a4740;letter-spacing:0.01em;font-size:13px;font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;">Refusals</th><td style="padding:9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;text-align:right;font-weight:700;color:#14130f;font-size:14px;font-family:ui-monospace,SFMono-Regular,Menlo,monospace;white-space:nowrap;">n/a</td></tr><tr><th scope="row" style="text-align:left;font-weight:600;padding:9px 8px 9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;color:#4a4740;letter-spacing:0.01em;font-size:13px;font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;">Context</th><td style="padding:9px 0;border:0;border-bottom:1px solid #cec9ba;background:none;text-align:right;font-weight:700;color:#14130f;font-size:14px;font-family:ui-monospace,SFMono-Regular,Menlo,monospace;white-space:nowrap;">256K</td></tr><tr><th scope="row" style="text-align:left;font-weight:600;padding:9px 8px 9px 0;border:0;background:none;color:#4a4740;letter-spacing:0.01em;font-size:13px;font-family:ui-sans-serif,system-ui,-apple-system,sans-serif;">MTP head</th><td style="padding:9px 0;border:0;background:none;text-align:right;font-weight:700;color:#b5231c;font-size:14px;font-family:ui-monospace,SFMono-Regular,Menlo,monospace;white-space:nowrap;">bf16</td></tr></table></div>

TL;DR: DavidAU's Qwen3.6-27B Fable-Fusion-711 (Uncensored/Heretic, with an MTP head), quantized to NVFP4 (W4A4) for vLLM on NVIDIA Blackwell. ~20 GB, 256K thinking VL, uncensored (from the base). The MTP head and vision tower are kept at bf16, 4-bit for the 27B backbone, full precision for speculative decoding and image input.

Qwen3.6-27B Fable-Fusion-711 NVFP4

DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP, quantized to NVFP4 (W4A4) in the compressed-tensors nvfp4-pack-quantized format with llm-compressor (GPTQ + MSE, shared fused-layer scales). This is a quantization only, the model is already uncensored/heretic upstream; no further abliteration was applied here.

  • —Built for vLLM on NVIDIA Blackwell (4-bit weight + 4-bit activation). Pre-Blackwell GPUs run it weight-only.
  • —The *MTP head (`mtp.) and vision tower (model.visual.`) are kept at bf16*, so speculative decoding and image input are unaffected by 4-bit.
  • —Loading and generation verified in vLLM on an NVIDIA GB10 (Blackwell, sm_121).
Uncensored model (inherited from the base). It follows instructions without refusal guardrails. You are responsible for how you use it.

Why the MTP head is kept at bf16

The base ships a multi-token prediction (MTP) head (mtp_num_hidden_layers: 1) that lets the model self-speculate: it drafts the next token(s) and verifies in one pass, cutting decode latency without a separate draft model.

That head sits on the acceptance-rate critical path. Quantizing it to 4-bit would push its drafts away from what the full-precision main model predicts, the verifier would reject more of them, and the speedup would shrink. So the recipe *ignores `re:.mtp.` and keeps the head at bf16 (it is tiny next to the 27B backbone). The 4-bit backbone does the heavy lifting; the full-precision MTP head keeps acceptance high. The vision tower (`re:.visual.*`) is kept at bf16 for the same fidelity reason.

Fidelity

Near-lossless versus the bf16 source, ~20 GB vs 55.6 GB bf16 (~36%). GPTQ error compensation and an MSE observer keep the drop from bf16 minimal; the header lists the full characteristics and Quantization covers the recipe.

Quickstart

NVFP4 is auto-detected from config.json (compressed-tensors); no quantization flag needed. --reasoning-parser qwen3 splits the <think> block into reasoning_content.

bash
vllm serve maci0/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4 \
  --served-model-name fable-fusion-27b-nvfp4 \
  --max-model-len 131072 \
  --gpu-memory-utilization 0.90 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3
  • —Supports up to 262144 tokens; keep at least 128K to preserve thinking quality.
  • —Add --language-model-only to skip the vision tower and free KV cache for text use.

Speculative decoding (MTP)

The bf16 MTP head enables vLLM's built-in self-speculative decoding:

bash
vllm serve maci0/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-NVFP4 \
  --served-model-name fable-fusion-27b-nvfp4 \
  --max-model-len 131072 --kv-cache-dtype fp8 --reasoning-parser qwen3 \
  --speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'

num_speculative_tokens matches the head depth (mtp_num_hidden_layers: 1). Acceptance is not penalized by quantization because the head is bf16.

Quantization

SchemeNVFP4, W4A4
Weight roundingGPTQ (Hessian-based error compensation), MSE observer
WeightsFP4 (E2M1), group_size=16, tensor_group, FP8 (E4M3) group scales, shared across fused layers
ActivationsFP4, dynamic per-group, FP8 (E4M3) scales
Quantizedall language-model Linear layers
Kept in bf16MTP head (mtp.*), vision tower (model.visual.*), lm_head
Untouchedgated delta-net Conv1d and SSM params (A_log, dt_bias), never Linear

The model is loaded with AutoModelForImageTextToText so the vision tower is quantized-and-kept (bf16) rather than dropped; the MTP head, which lives outside the transformers model graph, is re-attached at bf16 from the base. GPTQ is a quantization-time cost only; inference speed and format are identical to plain round-to-nearest NVFP4, but it chooses better 4-bit values.

About the base model

DavidAU's Fable-Fusion-711 is a 27B Qwen3.5-family (qwen3_5) vision-language model with thinking mode, a 256K context window, and an MTP head, already run through Heretic uncensoring.

  • —64 decoder layers; hybrid gated delta-net linear attention plus full attention; dense MLP; vision tower.
  • —MTP head (mtp_num_hidden_layers: 1) for speculative decoding.
  • —256K context (max_position_embeddings 262144).

Recommended sampling

Thinking mode is the default.

  • —Thinking, precise: temperature=0.6, top_p=0.95, top_k=20
  • —Thinking, general: temperature=1.0, top_p=0.95, top_k=20
  • —Instruct / non-thinking: temperature=0.7, top_p=0.80, top_k=20

Related

Notes

  • —Needs NVIDIA Blackwell (sm_121, e.g. GB10) for accelerated W4A4; pre-Blackwell GPUs run it weight-only.
  • —The MTP head and vision tower are bf16; enable speculative decoding via --speculative-config.
  • —--reasoning-parser is not auto-detected; pass it explicitly.
  • —No refusal guardrails; you are responsible for how you use it.

License

Apache-2.0, following the base model. Intended use and all responsibility for use follow the base model.

Credits

<div style="font-family:ui-monospace,SFMono-Regular,Menlo,monospace;font-size:12px;color:#6f6b60;border-top:1.5px solid #14130f;padding-top:14px;margin-top:30px;">Part of <a href="https://huggingface.co/spaces/maci0/rogue-quants" style="color:#b5231c;font-weight:700;text-decoration:none;">Rogue Quants</a> &middot; NVFP4 component datasheets &middot; <a href="https://huggingface.co/collections/maci0/nvfp4-quants-gb10-blackwell-6a446fc03174db196e436339" style="color:#b5231c;font-weight:700;text-decoration:none;">collection</a>. Fabricated on GB10 (Blackwell) with llm-compressor. Refusals shown per 100 harmful prompts; "n/a" = not separately measured (base-inherited).</div>