xero0000/Qwen3.8-27B-Palimpsest-GGUF
Qwen3.8-27B-Palimpsest-GGUF
GGUF releases of Qwen3.8-27B-Palimpsest, an experimental Qwen3.8-27B fine-tune for literary prose, continuity, structured tool use, and position-aware long-context behavior.
The flagship MIX-IQ3KT build is importance-calibrated and stays below the project's strict 12,000,000,000-byte language-model target. Conventional K-quant tiers are provided as controls and for broader runtime compatibility. The multimodal projector is not included in the language-model size budget.
Release status: evaluation in progress. Capability, MTP-equivalence, maximum-usable-context, and Forge gates are still running. Results below are local measurements with raw evidence; missing results are not treated as zero or silently inferred from the base model.
Fine-tune summary
Palimpsest merges two small LoRA stages into the original BF16 weights:
- a 1,152-example behavior/prose/tool mixture, selected at step 64; and
- an 80-step PoSE-style long-context stage updating Q/K LoRA modules only in the 16 full-attention layers, with virtual positions through 1M.
The full training recipe, source revisions, local development gates, and BF16 usage are documented in xero0000/Qwen3.8-27B-Palimpsest.
Quant files
The mixed file contains 866 tensors: 360 F32, 1 Q5K, 64 IQ4NL, 25 IQ3S, 176 IQ3KT, 144 IQ4KT, and 96 IQ3KS. It was quantized directly from the merged BF16 GGUF, never requantized from another lossy file.
Its importance matrix contains 433 entries from 256 chunks / 65,536 tokens balanced across tool use, prose, general text, and virtual-position examples. Calibration corpus SHA-256: 091f33bb6fb62d46850b3237cdc74b557344fe764adc53b31caab98808083bf1. Importance-matrix SHA-256: 1d6bb715ad4a36e756df9b30d059742ec809c50435107e6c2812408e3d154e6d.
Local quality screen
This deterministic 8+8 engineering screen is not a reproduction of the published Qwen model-card benchmark suite.
The three GPQA failures in both variants reached the output cap without the required final answer; none were server errors. Larger public benchmark runs and uncertainty-aware comparisons remain pending.
Controlled speed
Measured with expert-streaming build 516a0312 (build 4814), an RTX 3060 Ti 8 GB + RTX 2080 SUPER 8 GB + RTX 3080 10 GB, all 66 layers on GPU, 8/8/10 layer split, 24 threads, flash attention, batch 512, micro-batch 256, Q4_0 K/V cache, and three repetitions.
Q4KM is 4,790,167,552 bytes larger and decoded about 10.1% faster in the paired TG128 run. MIX-IQ3KT prefills 4096 tokens about 2.0% faster while remaining under 12 GB.
Measured cold-load/generation checks from the HDD were 19.15 seconds for Q4KM and 143.52 seconds for Q5KM. Storage cache state and quant kernels can strongly affect load time, so these are rig-specific measurements rather than universal expectations.
Tuned 512K allocation
The same three-GPU rig was also tuned with a 524,288-token allocation, static YaRN factor 2, CPU-resident Q4 KV, and a fixed 9,011-token prompt plus 128-token generation. This is a deployment-speed comparison at a 512K allocation, not a claim that a fully populated 512K prompt has passed the retrieval gate.
The tuned profile improved prompt throughput by 47.6% and decode throughput by 84.0%. MTP accepted 54 of 72 proposed tokens (75%) in its isolated controlled run. The 10/4/12 split deliberately reduces work on the RTX 2080 SUPER attached through the rig's PCIe x1 extender.
Runtime example
The calibrated mixed quant requires the included expert-streaming/IQK trellis runtime family. A tested native-context profile is:
llama-server \
-m Qwen3.8-27B-Palimpsest-MIX-IQ3KT.gguf \
--ctx-size 262144 --parallel 1 \
--n-gpu-layers 99 --split-mode layer --tensor-split 8,8,10 --main-gpu 2 \
--batch-size 512 --ubatch-size 256 --threads 24 --threads-batch 24 \
--flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --jinjaFor the tested 8 GB + 8 GB + 10 GB topology, the speed-tuned 512K profile adds:
--ctx-size 524288 --tensor-split 10,4,12 \
--batch-size 2048 --ubatch-size 512 --threads 8 --threads-batch 8 \
--cache-type-k q4_0 --cache-type-v q4_0 --no-kv-offload \
--rope-scaling yarn --rope-scale 2 --yarn-orig-ctx 262144 \
--spec-type mtp:n_max=1,p_min=0.0 --ctx-size-draft 8192Adjust the split for different cards. The conventional K-quants are intended to have broader llama.cpp-family compatibility, but Qwen3.8 architecture and MTP support still require a sufficiently recent build.
Context length
The underlying model is 262,144 tokens native. Do not use a 32,768-token YaRN origin with this model. For targets above 262,144, use a scale equal to target / 262144:
Static YaRN can reduce short-context quality, so enable it only for the target window you need. Measured 262K peaks leave insufficient aggregate headroom to double the Q4 KV cache safely, so this local 26 GB aggregate-VRAM rig uses CPU-resident Q4 KV (--no-kv-offload) beginning at 512K. A successful allocation alone is not evidence of usable context.
The release gate measures five-needle retrieval at approximately 1%, 25%, 50%, 75%, and 99%, exact instruction retention, prefill/decode latency, process RAM, and VRAM. The final card will report both the hard allocation maximum and the maximum window that passes the strict quality contract. It also reports a batch-usable maximum (strict quality, prefill no longer than 30 minutes, decode at least 5 tok/s) and an interactive-recommended maximum (strict quality, prefill no longer than 10 minutes, decode at least 10 tok/s). Peak process swap is recorded so a nominal pass cannot hide storage thrashing.
Measured native-context qualification
The flagship mixed quant passed all five retrieval positions and the exact one-line output contract at every native tier with context shifting disabled.
No model-process swap was observed. The proven interactive recommendation is 131,072 tokens; the proven native strict/batch maximum is 262,144 tokens. YaRN-scaled qualification remains in progress and is not inferred from these native results.
MTP / speculative decoding
The model retains one NextN/MTP layer. Initial single-run measurements on the flagship quant observed 24.62 target-only tok/s versus 31.96 tok/s with --spec-type mtp:n_max=1,p_min=0.0 (+29.81%, 112/142 drafted tokens accepted). The generated texts later diverged, so MTP remains opt-in until the repeated HTML workload and quality-equivalence gates finish.
Limitations
- Public capability and coding/agent benchmark results are not complete.
- Fully materialized long-context accuracy through 1M is still under test; the fine-tuning gate used short physical sequences with virtual positions.
- Vision weights were retained in BF16 training but were not fine-tuned or re-evaluated. A compatible projector is separate from these language GGUFs.
- English dominates the fine-tuning mixture.
- Generated facts, code, tool arguments, and safety-sensitive output require independent verification.
License
Apache-2.0, following the upstream Qwen3.8-27B release. Review upstream and dataset licenses before redistribution or commercial use.
