nuofang/Qwen3.8-9B-Distill-SLERP-F451-Pro-Writer-Uncensored-GGUF
<br>
这个仓库是 nuofang/Qwen3.8-9B-Distill-SLERP-F451-Pro-Writer-Uncensored 的 GGUF 量化模型
This repository contains GGUF quantization files for nuofang/Qwen3.8-9B-Distill-SLERP-F451-Pro-Writer-Uncensored.
<br>
imatrix 的校准数据以中文的小说、角色扮演为目标,同时保留逻辑和常识。仅在 Q5KM 及以下生效。
The calibration data for the imatrix is targeted at Chinese novels and role-playing (RP), while preserving logic and common sense. Effective only for Q5KM and lower.
<br><br>
如果你不知道下载哪种型号,我推荐 Q5_K_M,它通常与全精度相比感觉不到区别。手机或资源受限的cpu设备推荐IQ4XS。如果你需要mtp或现有的型号不能满足你,请查看 [mradermacher 的量化仓库](https://huggingface.co/mradermacher/Qwen3.8-9B-Distill-SLERP-F451-Pro-Writer-Uncensored-i1-GGUF) If you're unsure which quant to download, I recommend **Q5KM**—it generally feels indistinguishable from full precision. For mobile phones or resource-constrained CPU devices, IQ4XS is recommended. If you need MTP or if the existing quants don't meet your needs, check out mradermacher's quantization repo.
Perplexity Evaluation
(Tested against the provided calibration dataset)
- Base (F16/BF16): PPL = 14.1462 +/- 0.11497
- IQ4_XS: PPL = 12.0835 +/- 0.09560
- Q4_K_M: PPL = 12.0806 +/- 0.09554
- Q5_K_M: PPL = 12.0146 +/- 0.09496
- Q6_K: PPL = 11.9668 +/- 0.09463
- Q8_0: PPL = 11.9572 +/- 0.09459
If the perplexity drops after quantization compared to the original precision, it might not actually be an improvement. Instead, it could be caused by differences in how llama.cpp quantization and perplexity tools handle special tokens. 如果困惑度在量化之后与原精度相比变低,并不是真的提升,而是llamacpp量化工具和困惑度计算工具处理特殊token行为不同导致的。
