batiai/Qwen3.8-27B-GGUF
correct the calibration-corpus claim: it was English wikitext, not a multilingual mix
document tool calling: tags now advertise tools+thinking, with the token budget it needs
consistency pass on examples
use IQ4_XS consistently in the hero line and examples — q3 was the de-recommended tag
generalize the Q3_K finding across 4 model families and 2 chips; document the 30% num_ctx effect
explain the Metal speed picture: low-bit tiers are dequant-bound not bandwidth-bound, and point speed-first users at the MoE sibling
revise quant ladder from Apple Silicon measurements: IQ4_XS over Q3_K_M (smaller but slower on Metal)
rebuild complete: all six verified on llama.cpp and Ollama, sizes updated, num_ctx default documented
rebuild Qwen3.8-27B-Q6_K: prune blk.64 (MTP) + fix block_count metadata — fixes llama.cpp segfault and Ollama qwen3next init failure
rebuild Qwen3.8-27B-Q4_K_M: prune blk.64 (MTP) + fix block_count metadata — fixes llama.cpp segfault and Ollama qwen3next init failure
correct memory requirements (16GB does not fit), add Apple Silicon measurements and Ollama 0.20+ minimum
rebuild Qwen3.8-27B-IQ4_XS: prune blk.64 (MTP) + fix block_count metadata — fixes llama.cpp segfault and Ollama qwen3next init failure
rebuild Qwen3.8-27B-Q3_K_M: prune blk.64 (MTP) + fix block_count metadata — fixes llama.cpp segfault and Ollama qwen3next init failure
rebuild Qwen3.8-27B-IQ3_XXS: prune blk.64 (MTP) + fix block_count metadata — fixes llama.cpp segfault and Ollama qwen3next init failure
rebuild Qwen3.8-27B-Q2_K_S: prune blk.64 (MTP) + fix block_count metadata — fixes llama.cpp segfault and Ollama qwen3next init failure
disclose publish defect: Ollama tags fail to init, Q2/IQ3 segfault; all six rebuilding
withdraw Qwen3.8-27B-IQ3_XXS.gguf: --prune-layers 64 produced an unloadable file (segfault); rebuilding with blk.64 type override
withdraw Qwen3.8-27B-Q2_K_S.gguf: --prune-layers 64 produced an unloadable file (segfault); rebuilding with blk.64 type override
fix: nested code fence broke markdown rendering below the details block
Add files using upload-large-folder tool
docs: model overview, thinking-mode pitfall, measured speed
Upload README.md with huggingface_hub
initial commit
