protoLabsAI/ThinkingCap-Qwen3.6-27B-MTP-GGUF
Update README.md
Add IQ3_M-MTP (imatrix low-bit; quant_sensitivity 14/14, FC 51/54 — parity w/ NVFP4)
Add Q3_K_M-MTP (imatrix low-bit; quant_sensitivity 14/14, FC 51/54 — parity w/ NVFP4)
fix: set general.name for ThinkingCap-Qwen3.6-27B-Q8_0-MTP.gguf (was commit-hash)
fix: set general.name for ThinkingCap-Qwen3.6-27B-Q6_K-MTP.gguf (was commit-hash)
fix: set general.name for ThinkingCap-Qwen3.6-27B-Q5_K_M-MTP.gguf (was commit-hash)
fix: set general.name for ThinkingCap-Qwen3.6-27B-Q4_K_M-MTP.gguf (was commit-hash)
fix: set general.name for mtp-ThinkingCap-Qwen3.6-27B-head-Q8_0.gguf (was commit-hash)
fix: set general.name for mmproj-ThinkingCap-Qwen3.6-27B-f16.gguf (was commit-hash)
Add f16 vision mmproj (qwen3vl_merger, converted from base VLM tower; validated on trunk)
card: NVFP4 mini-ladder section (honest size curve)
NVFP4-Q4 mini-ladder rung (24GB dual-Blackwell target)
card: real benchmarks — ~40% brevity survives NVFP4, quant-sens 93%, FC 94%, MTP 62-91%
bf16-MTP
Q8_0-MTP
Q6_K-MTP
Q5_K_M-MTP
standalone MTP draft head
Q4_K_M-MTP
NVFP4-MTP (Blackwell FP4 + speculative decoding)
card: ThinkingCap NVFP4-MTP + GGUF ladder
initial commit
