AdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-GGUF
<img src="https://serve.thelabsource.com/u/6OiIHw.png" alt="adi-qwen3.5-4b-glm5.2-general" width="800">
adi-qwen3.5-4b-glm5.2-general
Part of the ADI (Advanced Data Intelligence) model line โ ADI Qwen3 series.
A small, fully local model that reasons and answers like a frontier teacher. Built by distilling glm-5.2 general-knowledge responses into a Qwen3.5-4B student with a bf16 LoRA fine-tune, then merged, converted, and quantized to GGUF. The student base retains native tool calling and a long context window.
Capabilities
Run it
Pull directly into Ollama:
ollama run hf.co/AdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-GGUF:Q4_K_MOr download the .gguf and point any llama.cpp-based runtime at it.
Try it live
A hosted demo is available as a Hugging Face Space โ chat with the model directly in your browser, no install required.
<a href="https://huggingface.co/spaces/AdvancedDataIntelligence/adi-qwen3.5-4b-glm5.2-general-demo"> <img src="https://serve.thelabsource.com/u/4Kb3iS.gif" alt="adi-qwen3.5-4b-glm5.2-general live demo" width="800"> </a>
Chat with the model directly in your browser โ no install required.
What this model is
This is a knowledge distillation: a strong teacher (glm-5.2) generated high-quality answers across ~2,000 diverse general-knowledge prompts, and the Qwen3.5-4B student was fine-tuned to imitate them. The result reasons and responds noticeably more like its teacher on general topics, while staying small enough to run on a single consumer GPU.
What distillation does โ and doesn't do. It transfers the teacher's reasoning style and answer quality, not net-new facts. A 4B model won't become an encyclopedia. For raw factual recall, retrieval-augmented generation (RAG) is the right tool, not fine-tuning. What you get here is a 4B that structures and explains like a much larger model on topics it already partly knows.
Training
The seed prompts were drawn from the human-written Databricks Dolly-15k dataset (filtered to remove items requiring an attached context passage, then deduplicated). The teacher was queried with thinking disabled so the student learns clean final answers rather than chain-of-thought it is too small to reproduce well.
Notes for re-builders
- Qwen3.5 trains in bf16 LoRA, not 4-bit QLoRA. Its gated-delta / Mamba-hybrid layers quantize poorly during training; 4-bit costs accuracy. bf16 LoRA uses ~10 GB on a 4B โ comfortable on a 16 GB card.
- Version pins: Qwen3.5 requires
transformers >= 5.2.0to be recognized, while the Unsloth training stack caps at<= 5.5.0. The working version istransformers == 5.5.0withnumpy < 2.3. - GGUF conversion was done with llama.cpp's
convert_hf_to_gguf.py, which already understands the Qwen3.5 SSM/MTP architecture.
Intended use
General-purpose local assistant: explanations, reasoning, Q&A, and tool-calling workflows where a small, private, offline-capable model is preferred over a hosted API. Not intended as a source of authoritative facts without retrieval.
License
Apache-2.0, inherited from the Qwen3.5-4B base model. You are free to use, modify, and redistribute under the terms of that license. Distilled training data was generated using glm-5.2; users should review the teacher model's terms for their own use case.
Built at [theLAB](https://thelabsource.com) โ Learning. Algorithms. Breakthroughs.
