jc-builds/Qwen3.5-0.8B-Q4_K_M-GGUF
3369
Qwen3.5-0.8B — GGUF (iPhone-optimized)
A Q4KM GGUF of `Qwen/Qwen3.5-0.8B` for on-device inference on any iPhone, iPad, or Apple Silicon Mac via llama.cpp or apps that wrap it (e.g. Haplo).
Hosted by jc-builds for the Haplo ecosystem. Quantization by Unsloth. Original weights © Alibaba Cloud, redistributed under the Apache 2.0 License.
TL;DR
The smallest model in Alibaba's Qwen3.5 family, using the same hybrid Gated DeltaNet + attention architecture as its larger siblings. At about half a gigabyte it runs on every supported device, and it is a large step up from the previous generation of sub-1B models. It runs in non-thinking mode by default.
Available quantizations
Details
How to use
Haplo (iPhone / iPad / Mac)
The model appears automatically in Haplo's model browser. Download URL:
https://huggingface.co/jc-builds/Qwen3.5-0.8B-Q4_K_M-GGUF/resolve/main/Qwen3.5-0.8B-Q4_K_M.ggufllama.cpp
llama-cli -hf jc-builds/Qwen3.5-0.8B-Q4_K_M-GGUF:Q4_K_MLicense
Apache 2.0. Qwen3.5 by Alibaba Cloud — see the upstream license.
