ahmedelkilani01/Swift-Qwen3.8-27B-i1-IQ4_XS-Smaller-GGUF
Swift-Qwen3.8-27B i1-IQ4_XS-Smaller GGUF
A 13.5 GB mixed 4-bit GGUF of UkisAI's Swift-Qwen3.8-27B, made for 16 GB GPUs. It keeps the MTP head and leaves room for a 64K q4_0 KV cache with MTP speculative decoding.
It uses the same per-tensor recipe as jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller, so the two files compare directly: same size and tensor types, different weights.
Why this quant
As of 2026-09-15, the published 4-bit Swift GGUFs are 15.3 GB or larger, which leaves too little of a 16 GB card for long context. Everything published between 12.5 and 14.8 GB is 3-bit (Q3KM, IQ3_XS and similar). This file keeps 4-bit attention, SSM and embedding tensors and takes the savings from the FFN.
Quick start
hf download ahmedelkilani01/Swift-Qwen3.8-27B-i1-IQ4_XS-Smaller-GGUF Swift-Qwen3.8-27b-i1-IQ4_XS-Smaller.gguf --local-dir .
llama-server -m Swift-Qwen3.8-27b-i1-IQ4_XS-Smaller.gguf \
-c 65536 -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1 --jinja \
--spec-type draft-mtp --spec-draft-n-max 2 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0- Sampling follows UkisAI's recommendation. The chat template accepts
reasoning_effortlow,mediumorxhigh(defaultxhigh) andenable_thinking. - MTP needs
--parallel 1. - For more context, drop MTP: about 104K fits on the same card without
--spec-type, at lower decode speed (measured on the jrell base quant, which has the identical layout).
Measured on a 16 GB card
AMD RX 6900 XT (16 GB, ROCm), llama.cpp build 10712 (daef7b687), with a desktop session running.
Evaluation
Every run compares this file with the jrell Qwen3.8 quant under identical settings on the card above. These are local single runs: treat small differences as noise, and don't compare them with published full-precision scores.
- IFEval: thinking on used temperature 0.7 and
--reasoning-budget 8192; thinking off was greedy. Neither difference is significant (McNemar p = 1.0 and 0.84). - Terminal-Bench, terminus-2: 5 easy, 5 medium and 5 hard tasks that have published per-task terminus-2 results; default timeouts, thinking
xhigh, temperature 1.0. Swift passed every task the base quant passed, plus two where the base timed out. Only 2 tasks differ, so the gap is not statistically significant (p = 0.5), but the lower agent time and token counts are what the fine-tune is meant to deliver. - Terminal-Bench, pi harness: the same tasks run through the pi coding agent 0.85.1 with settings you would actually use rather than the benchmark's defaults: 64K context with MTP, thinking
xhigh, native tool calls, 3× agent timeouts and a 30-minute request timeout. Both models score far higher than under terminus-2, mostly because long single responses no longer hit a 600 s client timeout. Swift leads again, on 3 discordant tasks (p = 1.0). Neither number is comparable to published leaderboard scores. - Long-context retrieval: thinking off, q4_0 KV,
-c 106496, no MTP; both models also got 5/5 single-key lookups at every depth.
How it was made
Requantized from mradermacher's Q8_0 (SHA-256 7dd7cc390443ab8a48ecddb216fa05bc087f6f9023c103aa13e4f6b9c61a285f) with mradermacher's imatrix (319 chunks × 512 tokens), using llama.cpp build 10712.
llama-quantize --allow-requantize --imatrix Swift-Qwen3.8-27b.imatrix.gguf \
--token-embedding-type iq4_xs --output-tensor-type q6_k \
--tensor-type 'ffn_(up|gate|down)\.weight=iq3_s' \
--tensor-type 'attn_(qkv|v)\.weight=q5_k' \
--tensor-type 'attn_(q|k|output|gate)\.weight=iq4_xs' \
--tensor-type 'ssm_(alpha|beta|out)\.weight=iq4_xs' \
--tensor-type 'nextn\.eh_proj\.weight=iq4_xs' \
Swift-Qwen3.8-27b.Q8_0.gguf Swift-Qwen3.8-27b-i1-IQ4_XS-Smaller.gguf IQ4_XS--tensor-typepatterns are regular expressions and the first match wins, hence the anchors.- The output was checked against the jrell file: identical tensor types in every tensor group across all 866 tensors,
block_count65 andnextn_predict_layers1 (MTP head present). - The imatrix has no entries for the MTP block, so those tensors were quantized without one.
- Requantizing from Q8_0 instead of BF16 adds a small extra loss; the perplexity match above suggests it is negligible.
Limitations
- All numbers come from single local runs on one card, and the Terminal-Bench subset is small.
- UkisAI's card reports small regressions against the base model on some benchmarks (for example AIME). Those were not re-tested here.
License
Swift is distributed under the Swift Open License v1.0. According to UkisAI's model card, personal, research, educational, evaluation and commercial use are free for individuals and organizations with annual recurring revenue (including affiliates) of up to US$1,000,000; above that, commercial use requires a Swift Enterprise License from UkisAI. The source repository does not include the full license text, so see the Swift model card and contact UkisAI for the exact terms. Swift is a derivative of Qwen/Qwen3.8-27B; check its license as well.
Credits
- UkisAI for Swift-Qwen3.8-27B
- The Qwen team for Qwen3.8-27B
- jrell for the IQ4_XS-Smaller recipe
- mradermacher for the Q8_0 source and the imatrix
- llama.cpp
