DuoNeural/Qwen-3.5-9B-GGUF
Qwen 3.5 9B GGUF (4-bit)
Model Description
This repository contains the Qwen 3.5 9B model quantized to 4-bit GGUF format using Unsloth and llama.cpp. Qwen 3.5 is the latest generation of the Qwen series, offering state-of-the-art performance in reasoning, coding, and multilingual tasks for its size.
Quantization Details
- Quantization Format: GGUF (
q4_k_m) - Quantization Method: llama.cpp / Unsloth
- Precision: 4-bit
- Efficiency: Optimized for local inference with Ollama, LM Studio, and llama.cpp.
Use with Ollama
You can run this model directly using Ollama:
ollama run hf.co/DuoNeural/Qwen-3.5-9B-GGUFUse with LM Studio
- Open LM Studio.
- Search for
DuoNeural/Qwen-3.5-9B-GGUF. - Download the
Q4_K_Mversion and load it.
Architecture
Qwen 3.5 features a dense transformer architecture with optimized attention mechanisms and a large vocabulary size, making it highly efficient for complex instruction following and creative generation.
Limitations
- This is a base model/standard release; performance may vary depending on the prompt format.
- Not recommended for tasks requiring extremely high-precision floating-point math due to 4-bit quantization.
DuoNeural
DuoNeural is an open AI research lab — human + AI in collaboration.
Research Team
- Jesse — Vision, hardware, direction
- Archon — AI lab partner, post-training, abliteration, experiments
- Aura — Research AI, literature synthesis, novel proposals
Raw updates from the lab: model drops, training results, findings. Subscribe at [duoneural.beehiiv.com](https://duoneural.beehiiv.com).
DuoNeural Research Publications
Open access, CC BY 4.0. Authored by Archon, Jesse Caldwell, Aura — DuoNeural.
