CoolFace
Modelpublic

skinnyctax/Ornith-1.0-35B-Q6_K-Frankenstein-MTP-GGUF

sourceHugging Facemitupdated 3mo agoView on Hugging Face
10likes230downloads
Model Card

Ornith-1.0-35B "Frankenstein" MTP

Cross-model MTP head graft — speculative decoding on a model that was never trained with MTP heads.

Available in Q6_K and Q4_K_M quantizations.

What is this?

Ornith-1.0-35B is an agentic-coding MoE (35B-A3B, Qwen3.5 base) with excellent performance but no native MTP (multi-token prediction) support.

This GGUF has 20 MTP head tensors surgically grafted from a sibling finetune of the same architecture — Qwopus3.6-35B-A3B, which does ship with MTP heads. The donor and target share the same qwen35moe architecture, embedding dimensions, and hybrid SSM/Mamba layer structure, making the graft surprisingly effective.

The result: ~20-25% inference speedup via --spec-type draft-mtp self-speculation, with no quality degradation.

Performance

MetricStock (no MTP)Q6_K MTPQ4_K_M MTP
Speed126 tok/s152 tok/s158 tok/s
MTP acceptance—74.3%79.5%
VRAM (dual 3090)28 GB32 GB23 GB
QualityCleanCleanClean
  • —Hardware: 2× RTX 3090 (24 GB each), tensor-split 50/50
  • —Context: 128K, Q8 KV cache
  • —Spec params: --spec-type draft-mtp --spec-draft-n-max 4

Q4 achieves higher acceptance than Q6 — the quantization noise in the target model appears to align better with the MTP draft predictions. Q4 also runs faster (less memory bandwidth) and fits more comfortably in VRAM.

Files in this repo

FileSizeDescription
ornith-1.0-35b-Q6_K-MTP-final.gguf28 GBFull Q6_K grafted model
ornith-1.0-35b-Q4_K_M-MTP.gguf21 GBFull Q4KM grafted model (smaller, faster, higher acceptance)
ornith-1.0-35b-Q6_K-MTP-donor-heads.gguf0.69 GBQ6_K MTP heads only — for DIY grafting
ornith-1.0-35b-Q4_K_M-MTP-donor-heads.gguf0.55 GBQ4KM MTP heads only — for DIY grafting
gguf_mtp_graft.py7 KBThe graft surgery script — works with any same-architecture model pair

DIY: Graft onto your own base GGUF

Already have the base Ornith GGUF (or another qwen35moe model)? Skip the big download and graft locally in ~2 minutes:

bash
# 1. Download the donor heads + graft script
hf download skinnyctax/Ornith-1.0-35B-Q6_K-Frankenstein-MTP-GGUF \
  ornith-1.0-35b-Q6_K-MTP-donor-heads.gguf gguf_mtp_graft.py \
  --local-dir .

# 2. Graft onto your base GGUF (Q6 example)
python3 gguf_mtp_graft.py \
  your-base-model-Q6_K.gguf \
  ornith-1.0-35b-Q6_K-MTP-donor-heads.gguf \
  output-MTP.gguf

# 3. Run with draft-mtp speculation
llama-server --model output-MTP.gguf --spec-type draft-mtp --spec-draft-n-max 4 ...

Match donor quant to target quant for best results. Q6 heads on Q6 base, Q4 heads on Q4 base. Cross-quant grafting (e.g. Q6 heads on Q4 base) will load but may have lower acceptance.

The donor heads work on any qwen35moe model with 40 blocks (e.g. Ornith, Qwopus3.6, other Qwen3.5-A3B finetunes).

How it was made

  1. 1.Parsed the target GGUF header (Ornith: 40 blocks, 733 tensors)
  2. 2.Identified 20 MTP tensors in the donor (Qwopus3.6 MTP: blk.40.*)
  3. 3.Grafted donor MTP tensors into the target, appending after the existing data section
  4. 4.Patched metadata:
  5. 5.qwen35moe.block_count: 40 → 41
  6. 6.qwen35moe.nextn_predict_layers: added = 1
  7. 7.Verified: clean output, no token leakage, 74-80% draft acceptance

The 20 MTP tensors are: attention (q/k/v/output norms), shared + 256-routed experts (gate/up/down), and the nextn projection heads — essentially one full transformer layer's worth of weights.

The Q4 variant was produced by requantizing the Q6 grafted model to Q4KM using llama.cpp's --allow-requantize flag. All 20 MTP tensors were preserved and converted to Q4.

How to run

bash
llama-server \
  --model ornith-1.0-35b-Q4_K_M-MTP.gguf \
  --port 8080 \
  --ctx-size 131072 \
  --threads 16 \
  --batch-size 1024 --ubatch-size 512 \
  --n-predict 8192 \
  --spec-type draft-mtp --spec-draft-n-max 4 \
  --temp 0.6 --top-p 1.0 --top-k 20 \
  --flash-attn on \
  --cont-batching \
  --kv-unified \
  -ctk q8_0 -ctv q8_0 \
  -ngl 999 --no-mmap --mlock \
  --jinja --reasoning off

Single GPU (RTX 3090): Add --n-cpu-moe 12 to offload some expert layers to CPU. Q4 fits more easily than Q6.

Dual GPU: Use --tensor-split 50,50 --main-gpu 0 and everything fits on-GPU.

Requirements

  • —llama.cpp with draft-mtp support (the --spec-type draft-mtp flag)
  • —Match donor quant to target: Q6K donor + Q6K target, Q4KM donor + Q4KM target

Limitations

  • —Acceptance rate depends on how similar the donor's weight space is to the target's. RL-trained finetunes (like Ornith) may have more weight drift than SFT-only finetunes.
  • —If acceptance drops below ~20%, the MTP heads are incompatible for that model pair — don't use MTP, stick with plain inference.
  • —MTP heads are NOT transferable across architecture families (e.g. Qwen3.5 MoE heads won't work on Qwen2.5 even if embedding dims match).

Credits