CoolFace
Modelpublic

lmcoleman/Qwen3.6-35B-A3B-MagicQuant-GGUF

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes323downloads
Model Card

Qwen3.6-35B-A3B-MagicQuant-GGUF

Derivative of Qwen3.6-35B-A3B, quantized using MagicQuant hybrid evolutionary per-tensor search.

Sibling repo with AMD-native (ROCmFPX fork-only) builds: lmcoleman/Qwen3.6-35B-A3B-ROCmFPX-GGUF.

Base Model

This is a derivative of Qwen3.6-35B-A3B. All credit for the base model architecture and weights goes to the original authors. The base model's license applies to this derivative.

Quantization Method

Quantized using [MagicQuant](https://github.com/lucasmcoleman/MagicQuant) hybrid evolutionary per-tensor quantization, based on the methodology by [magiccodingman](https://github.com/magiccodingman/MagicQuant-Wiki):

  • —Tensors are classified into sensitivity groups (Embeddings, Head, Query, Key, Output, FFN Up/Down, MoE Experts, Router)
  • —An evolutionary search finds the optimal quantization type per group, balancing size vs. perplexity
  • —Q4/Q5/Q6 tier targets are produced with different size-quality tradeoffs
  • —Small-row tensors and sensitivity-critical layers (embeddings, output head, router) are kept at F32/F16/BF16
  • —This is NOT a uniform quantization -- each tensor group gets its own optimal type

GGUF Files

FileSizeQuant
Qwen3.6-35B-A3B-Q4_K_M.gguf21.7 GBQ4 hybrid
Qwen3.6-35B-A3B-Q5_K_M.gguf25.4 GBQ5 hybrid
Qwen3.6-35B-A3B-Q6_K.gguf29.1 GBQ6 hybrid

Usage

LM Studio

  1. 1.Download the GGUF file of your preferred quantization tier
  2. 2.Place it in your LM Studio models directory
  3. 3.Load the model in LM Studio -- it will auto-detect the chat template
  4. 4.The model supports the base model's full context length

llama.cpp

bash
# Interactive chat (--jinja uses the model's embedded chat template, not a hardcoded one)
llama-cli -m Qwen3.6-35B-A3B-Q5_K_M.gguf -c 8192 --jinja -cnv

# Single prompt
llama-cli -m Qwen3.6-35B-A3B-Q5_K_M.gguf -c 8192 -p "Your prompt here"

# Server mode
llama-server -m Qwen3.6-35B-A3B-Q5_K_M.gguf -c 8192 --port 8080 --jinja

Python (llama-cpp-python)

python
from llama_cpp import Llama

llm = Llama(model_path="./Qwen3.6-35B-A3B-Q5_K_M.gguf", n_ctx=8192)
output = llm.create_chat_completion(
    messages=[
        {"role": "user", "content": "Hello, how are you?"}
    ]
)
print(output["choices"][0]["message"]["content"])

Serving: MTP Speculative Decoding

This model includes MTP ("nextn") draft tensors, enabling self-speculative decoding -- measured ~1.6-1.9x faster generation with a ~95% first-token accept rate (no separate draft model needed; it drafts from itself):

bash
llama-server -m Qwen3.6-35B-A3B-Q4_K_M.gguf -c 8192 --port 8080 --host 127.0.0.1 -ngl 99 -md Qwen3.6-35B-A3B-Q4_K_M.gguf --spec-type draft-mtp -ctk q8_0 -ctv q8_0 -fa on

Memory cost: MTP needs its own draft context alongside the main context, so serving with it uses roughly 2x the model's memory compared to serving without `-md/--spec-type draft-mtp`.

Caveats

  • —The base model's license (apache-2.0) applies to all derivative files
  • —Quantization reduces precision -- verify outputs for your specific use case
  • —The hybrid quantization assigns different precision to different tensor groups, which means quality characteristics may differ from uniform quantizations

Limitations

  • —Quantized models may exhibit subtle differences from the full-precision fine-tune
  • —This model inherits any limitations and biases present in the base model

Generated with [MagicQuant](https://github.com/lucasmcoleman/MagicQuant)