AdvancedDataIntelligence/adi-qwen2.5-14b-glm5.2-general-GGUF
<p align="center"> <img src="https://serve.thelabsource.com/u/dDBDQ8.png" width="760" alt="adi-qwen2.5-14b-glm5.2-general"> </p>
adi-qwen2.5-14b-glm5.2-general
Part of the ADI (Advanced Data Intelligence) model line — ADI Qwen series.
A compact, fully local model that reasons and answers like a frontier teacher. Built by distilling glm-5.2 general-knowledge responses into a Qwen2.5-14B-Instruct student with a 4-bit QLoRA fine-tune, then merged, converted, and quantized to GGUF. The largest general ADI model to date — more parametric headroom than the 8B, still small enough to run on a single 16 GB consumer GPU. The student base retains native tool calling and a long context window.
Capabilities
Run it
Pull directly into Ollama:
ollama run hf.co/AdvancedDataIntelligence/adi-qwen2.5-14b-glm5.2-general-GGUF:Q4_K_MOr download the .gguf and point any llama.cpp-based runtime at it.
What this model is
This is a knowledge distillation: a strong teacher (glm-5.2) generated high-quality answers across a clean general-knowledge prompt set, and the Qwen2.5-14B-Instruct student was fine-tuned to imitate them. The result reasons and responds noticeably more like its teacher on general topics, with the most headroom of any general model in the ADI line, while still fitting on a single consumer GPU.
What distillation does — and doesn't do. It transfers the teacher's reasoning style and answer quality, not net-new facts. A 14B model carries more parametric knowledge than the smaller ADI students, but it still isn't an encyclopedia. For raw factual recall, retrieval-augmented generation (RAG) is the right tool, not fine-tuning. What you get here is a 14B that structures and explains like a much larger model on topics it already partly knows.
Training
The seed prompts were drawn from the human-written Databricks Dolly-15k dataset (filtered to remove items requiring an attached context passage, then deduplicated). The teacher was queried with thinking disabled so the student learns clean final answers rather than chain-of-thought.
Notes for re-builders
- 4-bit QLoRA via Unsloth with gradient checkpointing ("unsloth" mode), maxseqlength 2048, per-device batch 1 × grad-accum 8, pagedadamw8bit, LoRA targeting all attention + MLP projections. Peak VRAM held at 12.05 GB on a 16 GB card.
- GGUF conversion was done via streaming LoRA merge → f16 GGUF (28 GB intermediate) → Q4KM quantize (8.4 GB, 4.87 bpw) with llama.cpp.
Intended use
General-purpose local assistant: explanations, reasoning, Q&A, and tool-calling workflows where a capable, private, offline-capable model is preferred over a hosted API. Not intended as a source of authoritative facts without retrieval.
License
Apache-2.0, inherited from the Qwen2.5-14B-Instruct base model. You are free to use, modify, and redistribute under the terms of that license. Distilled training data was generated using glm-5.2; users should review the teacher model's terms for their own use case.
Built at [theLAB](https://thelabsource.com) — Learning. Algorithms. Breakthroughs.
