devan-carlin/Qwen3.8-Flash-Next-W4A16
11.4k
Reword serving tip: describe failure mode as repetition/loops, not language
Document temperature guidance: run at 0.7, not the recommended 1.0; pin via override-generation-config
Note served arch is Qwen4ExpForConditionalGeneration (multimodal)
Revert text-only note: repo ships model.visual.* weights (333 tensors), model is multimodal
Fix model card: text-only (no visual.* weights shipped), correct MTP note
Upload README.md with huggingface_hub
Add real links: fork branch xpu-qwen4exp, setup script, patch, ops guide
Remove owner-only publishing section from model card
Polish model card language (Gemma4-assisted edit)
Qwen3.8-Flash-Next W4A16 (int4 g128) + PLE table, XPU serving recipe
initial commit
