Frosty40/SIQ-1-35B-B70-Turbo-GGUF
SIQ-1-35B · B70 Turbo · GGUF (Q5KM)
<p align="center"> <img src="hero.jpeg" alt="SIQ-1-35B · B70 Turbo · Q5KM GGUF" width="460"> </p>
Imatrix-quantized GGUF build of [`AlexWortega/SIQ-1-35B`](https://huggingface.co/AlexWortega/SIQ-1-35B) — a reasoning/agentic finetune of Qwen3.6-35B-A3B (MoE, 35B total / ~3B active, 256 experts, Qwen3 <think> template). Quantized with an importance matrix and tuned for fast local inference on Intel Arc.
🙌 The model is [AlexWortega](https://huggingface.co/AlexWortega)'s — [`AlexWortega/SIQ-1-35B`](https://huggingface.co/AlexWortega/SIQ-1-35B). This repo is only a fast GGUF build of it. The reasoning, the agentic behavior, all of it is his post-training. If you use this, go ★ the original — it deserves the reach.
This release ships `Q5_K_M` — the quant that holds all measurable reasoning (20/20 gate) while full-offloading on a 30 GiB B70. Q4 / Q6 / Q8 builds are evaluated and will be added to this repo later.
Files
Performance (Intel Arc B70, 30 GiB · SYCL — decode-fusion build)
Measured on the B70: flash-attention on, q8_0 KV cache, full offload (ngl 99).
- +9.8% from the Turbo fusions: decode 85.2 → 93.6 t/s vs stock llama.cpp SYCL, same GGUF weights (fusions on/off, 20/20 reasoning gate held). This is the controlled engine win.
- Trained context 256K; decode holds at depth (81 t/s @16K).
Turbo upgrades (vs stock llama.cpp SYCL)
- Prefill +53% —
-ub 4096un-starves the MoE expert GEMMs: 874 → 1334 t/s at 131072 ctx / q8_0 KV (20/20 reasoning gate held). +73% at small-ctx / f16-KV bench. - Deterministic decode, free —
GGML_SYCL_DISABLE_DNN=1; the native GEMM matches/beats oneDNN for this MoE, decode unchanged. - Decode +9.8% — MoE kernel fusions (expert-combine reduce + topk-moe router + gated-delta-net glue), same weights.
- Serve config:
GGML_SYCL_DISABLE_DNN=1 … -ngl 99 -fa on -ctk q8_0 -ctv q8_0 -c 131072 -b 8192 -ub 4096 --jinja
<sub>For reference the same-arch base Qwen3.6-35B-A3B runs ~75.6 t/s here, but it ships Q5_K_XL; that gap is mostly quant format (Q5KM is leaner), not optimization — the Turbo claim above is the same-weights fusion delta.</sub>
What "Turbo" buys (vs. stock llama.cpp SYCL): decode-path kernel fusions for this MoE — expert-combine reduce + topk-moe router fusion + gated-delta-net gate glue (see the companion repo). These collapse the per-layer tiny-op serial chains that bottleneck single-stream decode on this backend. Q5_K_M decode 85.2 → 93.6 t/s (+9.8%), correctness-neutral (20/20 reasoning gate held with fusions on vs off). It is a runtime/engine speedup — same GGUF weights, built with the patches applied; no re-quantization.
<sub>¹ wikitext-2, 100 chunks, ctx512. ² 20 verifiable multi-step problems graded vs a hand-verified key.</sub>
Why Q5KM
Upstream documents a GPQA-Diamond reasoning drop at Q4 (90.2 → 70.2). On the 20-problem verifiable gate, Q5/Q6/Q8 hold 20/20 while raw Q4 drops one probe (19/20) — and perplexity (Δ1.3% Q4→Q6) can't see that collapse, so the shipping pick comes from the reasoning gate, not PPL. Q5KM is the smallest quant that keeps full measurable reasoning and full-offloads on the 30 GiB B70.
Run (llama.cpp)
llama-server -m siq-1-35b-Q5_K_M.gguf \
-ngl 99 -c 32768 -fa on \
--cache-type-k q8_0 --cache-type-v q8_0 \
--jinja --host 0.0.0.0 --port 8092Use greedy (`temperature: 0`) for reasoning; the model thinks inside <think>…</think>, so allow a generous max_tokens (~3k on hard questions).
Credits & lineage
This package is a fast wrapper. The work that matters is theirs:
- ★ [AlexWortega](https://huggingface.co/AlexWortega) — [`SIQ-1-35B`](https://huggingface.co/AlexWortega/SIQ-1-35B) — the model. AlexWortega did the reasoning/agentic post-training that makes this thing actually think: every
<think>trace and every solved multi-step problem in the eval here is his finetune, not ours. It's under-followed for how good it is — [give the original a like](https://huggingface.co/AlexWortega/SIQ-1-35B). - [Qwen](https://huggingface.co/Qwen) (Alibaba) — `Qwen3.6-35B-A3B` — the base architecture SIQ builds on: the 35B / ~3B-active MoE, 256 experts, gated-delta-net attention, 256K context, and the
<think>template. - [llama.cpp / ggml](https://github.com/ggml-org/llama.cpp) (Georgi Gerganov & contributors) — the inference engine, the GGUF format, imatrix quantization, and the SYCL backend every optimization here plugs into.
- [Intel](https://github.com/intel/llvm) — oneAPI / SYCL and the Arc hardware this is tuned for.
GGUF quantization + B70 Turbo tuning (the only part that's ours): Frosty40. Everything above is the foundation; this repo just made it faster on one GPU. License Apache-2.0, inherited from the base.
