Flexan/kshitijthakkar-qwen3.5-moe-0.87B-d0.8B-GGUF
GGUF Files for qwen3.5-moe-0.87B-d0.8B
These are the GGUF files for kshitijthakkar/qwen3.5-moe-0.87B-d0.8B.
[!WARNING] This GGUF seems to be broken. This is according to my own testing. Feel free to try it out yourself, and please confirm or deny whether it's broken in the Community tab.
Downloads
Note from Flexan
I provide GGUFs and quantizations of publicly available models that do not have a GGUF equivalent available yet, usually for models I deem interesting and wish to try out.
If there are some quants missing that you'd like me to add, you may request one in the community tab. If you want to request a public model to be converted, you can also request that in the community tab. If you have questions regarding this model, please refer to the original model repo.
You can find more info about me and what I do here.
Qwen3.5 MoE 0.85B (from Qwen3.5-0.8B)
A Qwen3.5 Mixture-of-Experts model created via dual-source weight transfer:
- Backbone (attention, embeddings, vision, norms): from Qwen/Qwen3.5-0.8B
- MoE experts (routed + shared): from Qwen/Qwen3.5-35B-A3B (sliced 256->8 experts, bilinear resized)
Model Details
Design
Total MoE FFN parameters are approximately equal to the dense model's FFN parameters. The speed benefit comes from sparsity: only top-2 experts
- shared expert are active per token (~1/3 of total FFN).
Most weights are pre-trained (backbone from dense model, experts from 35B-A3B). Only the MoE dimension resize introduces noise, making this model suitable for fine-tuning at nominal cost.
Weight Transfer Sources
License
Apache 2.0 (following source models)
