ticeclock/Swift-Qwen3.8-27B-RCO-GGUF
233.8k
ETA Uploaded IQ3_S version, as well as the Swift mmproj so you don't have to go hunting for it.
A quant of ukisai/Swift-Qwen3.8-27B-GGUF naïvely quantized with tensor types from ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF. Usable with 16GB GPUs. Both quants here include the MTP head, I just named one file poorly. Tensor type file and imatrix are included for those who want to reproduce or improve upon this work.
All credit for the base model to Qwen team and UkisAI. All credit for the quant calculating to ISTA-DASLab. Two great tastes that taste great together!
Quantized with llama.cpp tag b11054. Quant command used:
llama-quantize \
--output-tensor-type IQ4_XS \
--token-embedding-type IQ2_S \
--imatrix imatrix-qwen3.8-27b.gguf \
--tensor-type-file tensor-types.txt \
Swift-Qwen3.8-27B-F16-00001-of-00003.gguf \
Swift-Qwen3.8-27B-RCO-IQ3_XXS.gguf \
IQ3_XXS