philbert440/Qwen3.8-27B-W4A16-AWQ
29497k
Fix tokenizer: drop leftover truncation block, restore upstream tokenizer.json/tokenizer_config.json (fixes vision >1024 tokens; see discussion #3)
Measured performance v2: warm multi-pass methodology, draft-mode comparison, instruct + concurrency regimes
Update README.md
Model card v3: base-model benchmarks, official sampling params, YaRN long-context, citation
Model card v2: badges, V100 throughput chart, full validation matrix, serving recipes
Model card: recipe, thinking-mode calibration, V100 validation matrix w/ MTP
W4A16-AWQ g128 asym mse, June-proven hybrid smoothing, Magpie thinking calib, bf16 MTP graft
initial commit
