TitanPythons/Qwen3.5-VL-9B-8bit-MLX-CRACK
Add Ko-fi support section
Add vMLX promotional banner to README
Add vMLX promotional banner to README
Upload model.safetensors.index.json with huggingface_hub
Upload tokenizer_config.json with huggingface_hub
Update: increase repetition_penalty to 1.3, add presence_penalty 1.5 per Qwen recommendation
Fix: disable thinking by default to prevent repetition loops on small models
Fix: disable thinking by default to prevent repetition loops on small models
Fix enable_thinking in tokenizer_config: match Qwen original (thinking ON by default for 4B+)
Fix enable_thinking: match Qwen original (thinking ON by default for 4B+)
Fix: inject chat_template with enable_thinking support into tokenizer_config.json
Fix standalone/spec-decoding positioning in README
Fix README: clarify standalone vs spec-decode role, remove redundant phrasing
Fix: remove incorrect draft model references (9B is standalone)
Fix: default thinking OFF to prevent CoT loop on CRACK models (F13-DENSE)
Upload generation_config.json with huggingface_hub
Upload config.json with huggingface_hub
Upload folder using huggingface_hub
initial commit
