CoolFace
Modelpublic

ahmedelkilani01/Swift-Qwen3.8-27B-i1-IQ4_XS-Smaller-GGUF

sourceHugging Faceotherupdated 12d agoView on Hugging Face
1likes448downloads
Model Card

Swift-Qwen3.8-27B i1-IQ4_XS-Smaller GGUF

A 13.5 GB mixed 4-bit GGUF of UkisAI's Swift-Qwen3.8-27B, made for 16 GB GPUs. It keeps the MTP head and leaves room for a 64K q4_0 KV cache with MTP speculative decoding.

It uses the same per-tensor recipe as jrell/Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller, so the two files compare directly: same size and tensor types, different weights.

FileSizeBPWSHA-256
Swift-Qwen3.8-27b-i1-IQ4_XS-Smaller.gguf13,543,868,896 B (12.61 GiB)3.96341492f26d8a3326114f4523b48af0aeb5035f4e6a16dcd86b14938b6e13bdec

Why this quant

As of 2026-09-15, the published 4-bit Swift GGUFs are 15.3 GB or larger, which leaves too little of a 16 GB card for long context. Everything published between 12.5 and 14.8 GB is 3-bit (Q3KM, IQ3_XS and similar). This file keeps 4-bit attention, SSM and embedding tensors and takes the savings from the FFN.

Quick start

bash
hf download ahmedelkilani01/Swift-Qwen3.8-27B-i1-IQ4_XS-Smaller-GGUF Swift-Qwen3.8-27b-i1-IQ4_XS-Smaller.gguf --local-dir .

llama-server -m Swift-Qwen3.8-27b-i1-IQ4_XS-Smaller.gguf \
  -c 65536 -ngl 999 -fa on -ctk q4_0 -ctv q4_0 --parallel 1 --jinja \
  --spec-type draft-mtp --spec-draft-n-max 2 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0
  • —Sampling follows UkisAI's recommendation. The chat template accepts reasoning_effort low, medium or xhigh (default xhigh) and enable_thinking.
  • —MTP needs --parallel 1.
  • —For more context, drop MTP: about 104K fits on the same card without --spec-type, at lower decode speed (measured on the jrell base quant, which has the identical layout).

Measured on a 16 GB card

AMD RX 6900 XT (16 GB, ROCm), llama.cpp build 10712 (daef7b687), with a desktop session running.

VRAM, 64K + MTP, q4_0 KV, during an agentic run14.6–14.9 GiB, fully on the GPU
Decode during Terminal-Bench (context up to ~58K, MTP)median 43.8 t/s (p10 39.8, p90 47.2), 116 requests
MTP draft acceptance, same run74.0% (jrell Qwen3.8 quant under the same settings: 69.4%)
Prompt processing, no MTP421 t/s at 33K, 363 t/s at 65K, 310 t/s at 101K

Evaluation

Every run compares this file with the jrell Qwen3.8 quant under identical settings on the card above. These are local single runs: treat small differences as noise, and don't compare them with published full-precision scores.

TestQwen3.8 base (jrell)Swift (this file)
Perplexity, 505 KB of llama.cpp docs, 2048 ctx × 8 chunks3.62663.6316
IFEval, thinking on, first 100 prompts, prompt-strict96.0%95.0%
IFEval, thinking off, first 270 prompts, prompt-strict81.5%82.2%
Long-context retrieval, 20 needles + 10 decoys, at 33K / 65K / 101K20/20 at each20/20 at each
Terminal-Bench 2.0, 15-task subset, terminus-2, 1 attempt60.0% (9/15)73.3% (11/15)
Same run: agent time / output tokens / timeouts12,604 s / 243K / 78,606 s / 188K / 4
Terminal-Bench 2.0, same subset, pi harness, real-use settings80.0% (12/15)86.7% (13/15)
Same run: agent time / output tokens / timeouts13,613 s / 399K / 211,437 s / 312K / 1
  • —IFEval: thinking on used temperature 0.7 and --reasoning-budget 8192; thinking off was greedy. Neither difference is significant (McNemar p = 1.0 and 0.84).
  • —Terminal-Bench, terminus-2: 5 easy, 5 medium and 5 hard tasks that have published per-task terminus-2 results; default timeouts, thinking xhigh, temperature 1.0. Swift passed every task the base quant passed, plus two where the base timed out. Only 2 tasks differ, so the gap is not statistically significant (p = 0.5), but the lower agent time and token counts are what the fine-tune is meant to deliver.
  • —Terminal-Bench, pi harness: the same tasks run through the pi coding agent 0.85.1 with settings you would actually use rather than the benchmark's defaults: 64K context with MTP, thinking xhigh, native tool calls, 3× agent timeouts and a 30-minute request timeout. Both models score far higher than under terminus-2, mostly because long single responses no longer hit a 600 s client timeout. Swift leads again, on 3 discordant tasks (p = 1.0). Neither number is comparable to published leaderboard scores.
  • —Long-context retrieval: thinking off, q4_0 KV, -c 106496, no MTP; both models also got 5/5 single-key lookups at every depth.

How it was made

Requantized from mradermacher's Q8_0 (SHA-256 7dd7cc390443ab8a48ecddb216fa05bc087f6f9023c103aa13e4f6b9c61a285f) with mradermacher's imatrix (319 chunks × 512 tokens), using llama.cpp build 10712.

TensorsType
token_embdIQ4_XS
outputQ6_K
ffn_up, ffn_gate, ffn_down (all 65 blocks)IQ3_S
attn_qkv (48 Gated DeltaNet blocks), attn_v (16 full-attention blocks + MTP block)Q5_K
attn_q, attn_k, attn_output, attn_gate, ssm_alpha, ssm_beta, ssm_out, nextn.eh_projIQ4_XS
Norms, ssm_a, ssm_conv1d, ssm_dtF32
bash
llama-quantize --allow-requantize --imatrix Swift-Qwen3.8-27b.imatrix.gguf \
  --token-embedding-type iq4_xs --output-tensor-type q6_k \
  --tensor-type 'ffn_(up|gate|down)\.weight=iq3_s' \
  --tensor-type 'attn_(qkv|v)\.weight=q5_k' \
  --tensor-type 'attn_(q|k|output|gate)\.weight=iq4_xs' \
  --tensor-type 'ssm_(alpha|beta|out)\.weight=iq4_xs' \
  --tensor-type 'nextn\.eh_proj\.weight=iq4_xs' \
  Swift-Qwen3.8-27b.Q8_0.gguf Swift-Qwen3.8-27b-i1-IQ4_XS-Smaller.gguf IQ4_XS
  • —--tensor-type patterns are regular expressions and the first match wins, hence the anchors.
  • —The output was checked against the jrell file: identical tensor types in every tensor group across all 866 tensors, block_count 65 and nextn_predict_layers 1 (MTP head present).
  • —The imatrix has no entries for the MTP block, so those tensors were quantized without one.
  • —Requantizing from Q8_0 instead of BF16 adds a small extra loss; the perplexity match above suggests it is negligible.

Limitations

  • —All numbers come from single local runs on one card, and the Terminal-Bench subset is small.
  • —UkisAI's card reports small regressions against the base model on some benchmarks (for example AIME). Those were not re-tested here.

License

Swift is distributed under the Swift Open License v1.0. According to UkisAI's model card, personal, research, educational, evaluation and commercial use are free for individuals and organizations with annual recurring revenue (including affiliates) of up to US$1,000,000; above that, commercial use requires a Swift Enterprise License from UkisAI. The source repository does not include the full license text, so see the Swift model card and contact UkisAI for the exact terms. Swift is a derivative of Qwen/Qwen3.8-27B; check its license as well.

Credits

  • —UkisAI for Swift-Qwen3.8-27B
  • —The Qwen team for Qwen3.8-27B
  • —jrell for the IQ4_XS-Smaller recipe
  • —mradermacher for the Q8_0 source and the imatrix
  • —llama.cpp