CoolFace
Modelpublic

Flexan/kshitijthakkar-qwen3.5-moe-0.87B-d0.8B-GGUF

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
2likes453downloads
Model Card

GGUF Files for qwen3.5-moe-0.87B-d0.8B

These are the GGUF files for kshitijthakkar/qwen3.5-moe-0.87B-d0.8B.

[!WARNING] This GGUF seems to be broken. This is according to my own testing. Feel free to try it out yourself, and please confirm or deny whether it's broken in the Community tab.

Downloads

GGUF LinkQuantizationDescription
DownloadQ2_KLowest quality
DownloadQ3KS
DownloadIQ3_SInteger quant, preferable over Q3KS
DownloadIQ3_MInteger quant
DownloadQ3KM
DownloadQ3KL
DownloadIQ4_XSInteger quant
DownloadQ4KSFast with good performance
DownloadQ4KMRecommended: Perfect mix of speed and performance
DownloadQ5KS
DownloadQ5KM
DownloadQ6_KVery good quality
Downloadf16Full precision, don't bother; use a quant

Note from Flexan

I provide GGUFs and quantizations of publicly available models that do not have a GGUF equivalent available yet, usually for models I deem interesting and wish to try out.

If there are some quants missing that you'd like me to add, you may request one in the community tab. If you want to request a public model to be converted, you can also request that in the community tab. If you have questions regarding this model, please refer to the original model repo.

You can find more info about me and what I do here.

Qwen3.5 MoE 0.85B (from Qwen3.5-0.8B)

A Qwen3.5 Mixture-of-Experts model created via dual-source weight transfer:

Model Details

PropertyValue
Total Parameters854,386,752 (0.85B)
Active Parameters677,439,552 (0.68B)
ArchitectureQwen3.5 Hybrid MoE
Experts8 routed + 1 shared, top-2
Hidden Size1024
Layers24 (hybrid: DeltaNet + full attention)
AttentionGQA 8Q / 2KV, head_dim=256
Context262,144 tokens
Vocab248,320
Dtypebfloat16

Design

Total MoE FFN parameters are approximately equal to the dense model's FFN parameters. The speed benefit comes from sparsity: only top-2 experts

  • —shared expert are active per token (~1/3 of total FFN).

Most weights are pre-trained (backbone from dense model, experts from 35B-A3B). Only the MoE dimension resize introduces noise, making this model suitable for fine-tuning at nominal cost.

Weight Transfer Sources

ComponentSourceStrategy
Embeddings, LM HeadQwen/Qwen3.5-0.8BExact copy
Attention (Q/K/V/O, norms)Qwen/Qwen3.5-0.8BExact copy
DeltaNet (linear attention)Qwen/Qwen3.5-0.8BExact copy
Vision encoderQwen/Qwen3.5-0.8BExact copy
Layer normsQwen/Qwen3.5-0.8BExact copy
Routed expertsQwen3.5-35B-A3BSlice 256->8, bilinear resize
Shared expertQwen3.5-35B-A3BBilinear resize
RouterQwen3.5-35B-A3BSlice + resize

License

Apache 2.0 (following source models)