goldhub/Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-Heretic-Uncensored-W4A16-AutoRound
Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-Heretic-Uncensored — W4A16 AutoRound
4-bit weight-only quantization (W4A16) of DavidAU's 40B "Chimera" — the 6-Core Fable Fusion Heretic Uncensored model. Quantized with Intel AutoRound recipe best, group_size 64, calibrated on high-quality reasoning traces.
Base model: DavidAU/Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-Heretic-Uncensored
Quantization Recipe
Unquantized layers (kept BF16): lm_head, embed_tokens, model.visual (vision tower), all linear_attn projections (Gated DeltaNet), and MTP head via WOQ[RTN].
Hybrid Architecture Note
This model uses a 3:1 hybrid attention pattern — 72 linear-attention layers (Gated DeltaNet with Mamba SSM) and 24 full-attention layers. During quantization, --bs 1 is mandatory: the partial rotary embedding (partial_rotary_factor=0.25, head_dim=256) causes batch-dimension mismatches at higher batch sizes on full-attention layers.
Benchmarks (Custom, RTX 3090)
Custom generation benchmark, temperature=0.6, vLLM TP=2.
Quality verdict: Reasoning task solved correctly (both the sheep logic puzzle and the two-train meeting problem). Code generation produced full production-ready Python with bitarray optimization, type hints, docstrings, error handling, and CLI benchmarking. Creative output sustained 8K+ tokens coherently.
Reference — base 27B Qwopus INT4 on same harness: 42–59 tok/s. The 40B quant is slower (more params) but produces longer, deeper reasoning chains.
Usage
vLLM (recommended, TP=2)
vllm serve goldhub/Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-Heretic-Uncensored-W4A16-AutoRound \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90 \
--trust-remote-codeSGLang (TP=1 PP=3)
This quantized Chimera successfully loads under SGLang with TP=1, PP=3:
python -m sglang.launch_server \
--model-path goldhub/Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-Heretic-Uncensored-W4A16-AutoRound \
--tp 1 --pp 3 \
--trust-remote-codeMTP (Multi-Token Prediction)
The MTP head is preserved (mtp.layers.0, mtp.fc) and quantized via WOQ[RTN]. Under vLLM with MTP=3, speculative decoding is functional with variable acceptance (observed 13%–100% depending on token predictability), yielding ~25–45 tok/s effective throughput. For maximum acceptance stability, MTP=2 is recommended for long-form generation.
SGLang + MTP caveat
On SGLang with --pp 3 (pipeline parallel), MTP speculative decoding does NOT work due to overlap between pipeline stages. Use --tp 2 or vLLM if you need MTP. On vLLM with --tensor-parallel-size 2, MTP=3 works and is confirmed alive (acceptance 70-100%).
Chat Template
This repo ships a patched chat template (chat_template.jinja) from Qwen-Fixed-Chat-Templates v21.3 (froggeric/Qwen-Fixed-Chat-Templates).
The original template is preserved as chat_template.original.jinja for reference. The patched template fixes tokenizer/chat-formatting issues and improves instruction-following fidelity.
Acknowledgments
- DavidAU — for the base 6-Core Chimera fine-tune
- Intel AutoRound team — quantization framework
- froggeric — fixed chat templates
- Quantized by Viktor Zhuromskyy on consumer hardware (RTX 3090)
License
This quantized model inherits the base model's license. AutoRound is Apache-2.0.
🚀 Long Story Short: W4G64 AutoRound Quantization & MTP Confirmation
Successfully quantized the 40B Chimera to INT4 (W4G64) using AutoRound.
- Weight: ~38 GB
- VRAM: Fits comfortably on multi-GPU setups (e.g., 2x24GB or 3xRTX3090/4090) with room for large context.
⚡ MTP (Multi-Token Prediction) Status: ALIVE & KICKING
Initial concerns about MTP breaking during quantization or expansion are debunked. Telemetry confirms MTP=3 is active and providing significant generation speedups without degrading the "Heretic" creative quality or reasoning capabilities.
📊 W4G64 Benchmark Highlights (vLLM / SGLang)
🛠️ How to run (vLLM example)
vllm serve goldhub/Qwen3.6-40B-Fable-Fusion-6-Core-Deckard-Eleanor-Heretic-Uncensored-W4A16-AutoRound \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--trust-remote-code