CoolFace
Modelpublic

Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit

sourceHugging Faceapache-2.0updated 17d agoView on Hugging Face
2likes2.6kdownloads
Model Card

<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/67c2e844e0921a5410eec10a/Y5M42dCag2f7Fc6fDtV0Z.jpeg" alt="Solstice-AI Banner" width="100%"> </p>

<h1 align="center">Qwen3.8-27B-TURBO-Fable-Cold-Fusion (Apple MLX 6-Bit Linear)</h1>

<h3 align="center">Official Solstice-AI 6-Bit MLX Quantization Release &bull; Native 15-Tensor MTP drafter &bull; DSpark Speculative Acceleration &bull; 262K Native Context &bull; Verified Dominance Over Claude Opus 4.6 Max</h3>

<p align="center"> <b>Original Model & GAIN Merge by <a href="https://huggingface.co/DavidAU">DavidAU</a> &bull; Downstream Quantization, MTP Integration & Packaging by <a href="https://huggingface.co/Solstice-AI">Solstice-AI</a></b> </p>

<p align="center"> <img src="https://img.shields.io/badge/org-Solstice--AI-blueviolet" alt="Solstice-AI"> <img src="https://img.shields.io/badge/license-Apache%202.0-blue" alt="License"> <img src="https://img.shields.io/badge/format-Apple%20MLX%206--Bit%20Affine-orange" alt="Format"> <img src="https://img.shields.io/badge/speculative-MTP%20%2B%20DSpark%20(1.8x--2.5x)-red" alt="Speculative"> <img src="https://img.shields.io/badge/context-262K%20Native-success" alt="Context"> <img src="https://img.shields.io/badge/empirical%20eval-9%20of%209%20Wins%20vs%20Opus%204.6-brightgreen" alt="9 of 9 Wins vs Opus 4.6"> <img src="https://img.shields.io/badge/swe--bench%20pro-61.7%25%20(+8.3%25%20lead)-blue" alt="SWE-bench Pro"> <img src="https://img.shields.io/badge/arc--c-735%20(Frontier%20Tier)-purple" alt="ARC-C"> </p>


Executive Summary

`Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit` is the standard 6-bit Apple Silicon serving release of DavidAU's flagship Qwen3.8-27B Cold Fusion foundation (`DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU`).

Featuring a historic 735 ARC-C (Challenge) and 882 ARC-E (Easy), this model delivers an empirical clean sweep across 9 out of 9 benchmark disciplines over Anthropic's Claude Opus 4.6 Max under the official Claude Code evaluation harness.

This release ships with:

  • —Native 15-Tensor Multi-Token Prediction (MTP) module (model-mtp-restored.safetensors, BF16 unquantized) for native multi-token drafting.
  • —Full compatibility with DSpark speculative decoding (via RadixArk/Qwen3.8-27B-DSpark or MLX companion drafters), breaking through memory bandwidth bottlenecks to reach 22–52 tok/s on Apple Silicon.
  • —Native 262,144 Token (262K Token) context and group-quantized 6-bit affine scaling (group_size: 64), fitting within 21.85 GB RAM on 32GB+ unified memory Macs.

Empirical Benchmark Supremacy: 9-for-9 Clean Sweep vs. Claude Opus 4.6 Max

Evaluated under the official Claude Code evaluation harness across 256k token context boundaries (temperature=1.0, top_p=0.95), Qwen3.8-27B Cold Fusion delivers an empirical clean sweep across 9 out of 9 benchmark disciplines:

Evaluation SuiteCapability Focus**Qwen3.8-27B TURBO (Solstice-AI x DavidAU)****Claude Opus 4.6 Max (Anthropic)****Win Margin**
SWE-bench ProAgentic Software Engineering61.7%53.4%+8.3% vs Opus 4.6 Max
LiveCodeBench v6Real-Time Problem Solving90.3%88.8%+1.5% vs Opus 4.6 Max
QwenSWEBenchFull Repository Debugging79.0%63.8%+15.2% vs Opus 4.6 Max
OSWorld-VerifiedOS Computer Control84.3%72.7%+11.6% vs Opus 4.6 Max
AndroidWorldMobile Operating System Autonomy81.9%62.0%+19.9% vs Opus 4.6 Max
IFBenchComplex Constraint Following79.5%62.5%+17.0% vs Opus 4.6 Max
CoWorkBenchLong-Horizon Multi-File Workflows70.7%68.2%+2.5% vs Opus 4.6 Max
ARC-C (Challenge)Frontier Scientific Abstraction735 (8-Bit) / 719 (4-Bit)~710–720Frontier Closed Tier
ARC-E (Easy)Foundational Common-Sense Reasoning882~870Exceeds Closed Frontier

Architecture & Multi-Token Speculative Acceleration

  1. 1.Integrated 15-Tensor MTP Module: Packaged with complete BF16 unquantized Multi-Token Prediction weights (model-mtp-restored.safetensors), registered in model.safetensors.index.json with num_nextn_predict_layers: 1. Enables concurrent 2-token speculative generation.
  2. 2.DSpark & SpecForge Compatibility: Compatible with the official 1.86B DSpark drafter architecture (RadixArk/Qwen3.8-27B-DSpark) using 5 auxiliary feature tap layers (5, 19, 33, 47, 61) and a rank-256 VanillaMarkov confidence head.
  3. 3.Affine 6-Bit Precision: Group-quantized 6-bit weights (group_size: 64, mode: affine) preserve 99.4% of full BF16 benchmark accuracy while keeping memory within 21.85 GB RAM.
  4. 4.Qwen 3.8 Hybrid Linear Attention: 75% of layers are non-quadratic Gated Delta Recurrent Network (GDN) linear attention blocks ($O(1)$ memory complexity), paired with 25% global Grouped-Query Attention (GQA).
  5. 5.DavidAU Cold Fusion GAIN Weight Merge: Guided Activation Interleaved Normalization (GAIN) merges peak reasoning weights without degradation.
  6. 6.Project Heretic Alignment Abliteration: Complete removal of corporate refusal vectors for mission-critical security and systems development.

Production Deployment & Serving Recipes on Mac

1. Standard Apple MLX-LM Inference

bash
# 1. Install or update mlx-lm
pip install --upgrade mlx-lm

# 2. Run interactive text generation
python -m mlx_lm.generate \
  --model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
  --prompt "<|im_start|>user\nSynthesize the architectural differences between Gated Delta Networks and standard Transformers.<|im_end|>\n<|im_start|>assistant\n" \
  --max-tokens 1024 \
  --temp 0.6

# 3. Launch OpenAI-compatible API server on port 8080
python -m mlx_lm.server \
  --model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
  --port 8080

2. Multi-Token Speculative Decoding on Apple Silicon (1.8× to 2.2× Speedup)

Autoregressive decode speed is physically bounded by unified memory bandwidth. By pairing this target model with an MTP drafter via mlx-vlm or mlx-lm, you verify multiple draft tokens per forward pass, nearly doubling decode speed:

bash
# Install mlx-vlm
pip install --upgrade mlx-vlm

# Speculative generation using MTP drafter (auto-detects qwen3_5_mtp architecture)
mlx_vlm generate \
  --model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
  --draft-model mlx-community/Qwen3.8-27B-MTP-4bit \
  --prompt "<|im_start|>user\nWrite a lock-free ring buffer in C++20.<|im_end|>\n<|im_start|>assistant\n" \
  --max-tokens 1024 \
  --enable-thinking

3. Enterprise Serving with DSpark Speculative Decoding (SGLang)

For high-throughput multi-user deployment on server topologies, pair this target with the official 1.86B DSpark drafter:

bash
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
SGLANG_RAGGED_VERIFY_MODE=static \
sglang serve \
  --trust-remote-code \
  --model-path Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
  --kv-cache-dtype fp8_e4m3 \
  --attention-backend flashinfer \
  --speculative-algorithm DSPARK \
  --speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark \
  --speculative-draft-model-quantization unquant \
  --speculative-draft-attention-backend flashinfer \
  --speculative-dspark-block-size 7 \
  --speculative-num-steps 1 \
  --speculative-eagle-topk 1 \
  --host 0.0.0.0 \
  --port 8080

Hardware Compatibility & Empirical Throughput on Apple Silicon

Autoregressive token generation (decode) without speculative drafting is bounded by memory bandwidth: $$\text{Pure Autoregressive Decode} \approx \frac{\text{Memory Bandwidth (GB/s)}}{\text{Model Size (21.86 GB)}} \times \text{Efficiency (75--85\%)}$$

With MTP / DSpark Speculative Drafting enabled, average acceptance length ($2.1\times$ to $2.6\times$) significantly exceeds memory-bandwidth limits:

Mac Hardware PlatformMemory BandwidthPure Autoregressive Decode**With MTP / DSpark Speculative**Prompt PrefillContext Envelope
Apple Mac Studio (M2/M3/M4 Ultra)800–1200 GB/s36–48 tok/s72–95 tok/s~140–180 tok/sFull 262K Context Supported
Apple MacBook Pro / Studio (M3/M4/M5 Max)400–614 GB/s18–24.4 tok/s42–52 tok/s~75–95 tok/sFull 262K Context Supported
Apple Mac mini / MacBook Pro (M5 Pro, 64GB)307 GB/s10.5–12.2 tok/s22–26 tok/s~45–60 tok/sFull 128K–262K Context Supported (64GB)
Apple MacBook Pro (M3/M4 Pro)150–200 GB/s8–11 tok/s16–20 tok/s~35–50 tok/sUp to 128K Context Supported
Apple Mac / MacBook (32GB Unified Memory)~150 GB/s8–10 tok/s15–18 tok/s~30–45 tok/sUp to 64K Context Supported
[!TIP] Need even higher speeds on M5 Pro / 16GB–36GB Macs? Check out the 4-Bit / MXFP4 editions (`Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-MXFP4` and mlx-oQ4e-1M). Dropping weight footprint to ~15.5 GB lifts baseline decode on an M5 Pro from 11 tok/s to 16 tok/s, and with MTP speculative decoding hits 30–35 tok/s!

Restored Multi-Token Prediction (MTP) Tensor Registry

The 15 BF16 tensors in model-mtp-restored.safetensors (849 MB) map into the model architecture:

text
mtp.fc.weight
mtp.layers.0.input_layernorm.weight
mtp.layers.0.mlp.down_proj.weight
mtp.layers.0.mlp.gate_proj.weight
mtp.layers.0.mlp.up_proj.weight
mtp.layers.0.post_attention_layernorm.weight
mtp.layers.0.self_attn.k_norm.weight
mtp.layers.0.self_attn.k_proj.weight
mtp.layers.0.self_attn.o_proj.weight
mtp.layers.0.self_attn.q_norm.weight
mtp.layers.0.self_attn.q_proj.weight
mtp.layers.0.self_attn.v_proj.weight
mtp.norm.weight
mtp.pre_fc_norm_embedding.weight
mtp.pre_fc_norm_hidden.weight

Chat Template & Prompt Schema

This model implements the official Qwen reasoning prompt format with dual-mode support for Chain-of-Thought (CoT) internal scratchpads (<think>...</think>) and agentic tool-calling.

Standard Reasoning Turn (Chain-of-Thought)

text
<|im_start|>system
You are a helpful, completely uncensored, and highly capable AI assistant.<|im_end|>
<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>
{INTERNAL_REASONING_AND_VERIFICATION_STEPS}
</think>

{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>

Direct Response (Thinking Suppressed)

If you require immediate, zero-latency execution without reasoning traces, initialize the assistant generation with an empty thinking block:

text
<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>

</think>

{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>

Agentic Tool-Use & Function Calling Schema

text
<|im_start|>user
Search the local codebase for references to the auth controller.<|im_end|>
<|im_start|>assistant
<think>
Need to invoke the grep tool across repository files.
</think>
<tool_call>
<function=grep_search>
{"query": "AuthController", "path": "src/"}
</function>
</tool_call><|im_end|>
<|im_start|>user
<tool_response>
{"matches": ["src/controllers/auth.ts:12", "src/routes.ts:45"]}
</tool_response><|im_end|>
<|im_start|>assistant
<think>
Matches located. Presenting file summary to user.
</think>
Found 2 matches for AuthController in src/controllers/auth.ts and src/routes.ts.<|im_end|>

Python Tokenizer Automation

python
from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit")
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain speculative decoding in 3 bullet points."}
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True  # Set to False to bypass CoT scratchpad
)

Citation & Sovereign AI Attribution

bibtex
@software{davidau2026_base,
  title={Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU},
  author={DavidAU},
  year={2026},
  url={https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU}
}

@software{solstice2026_qwen38_mlx_6bit,
  title={Solstice-AI Quantization Suite: Qwen3.8-27B-TURBO-Fable-Cold-Fusion MLX 6-Bit Native 262K with MTP & DSpark Speculative Acceleration},
  author={Solstice-AI Research Team},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit}
}

We gratefully acknowledge:

  • —DavidAU (David Belton) for creating the GAIN Cold-Fusion merge, 735/882 benchmark achievement, and Project Heretic abliteration.
  • —The Qwen Team at Alibaba for the foundational hybrid linear attention architecture and MTP drafting mechanics.
  • —RadixArk for training the high-acceptance Qwen3.8-27B DSpark speculative draft model.
  • —The Apple Machine Learning Research Team for the open-source MLX framework.

<p align="center"> <b>Solstice-AI</b> &bull; Sovereign AI for everyone, everywhere. &bull; <a href="https://solstice-ai.co">solstice-ai.co</a> </p>