nivvis/Qwen3.5-9B-EQ-v5.1-GGUF
Qwen3.5-9B-Heretic-v2-EQ-v5.1-GGUF
GGUF quantizations of nivvis/Qwen3.5-9B-EQ-v5.1 for llama.cpp, Ollama, and LM Studio.
Example outputs: EQ 9B v5.1 vs Vanilla Qwen3.5-9B — same prompt, same settings.
Available quantizations
Qwen3.5-9B-EQ-v5.1
A DPO fine-tune of trohrbaugh/Qwen3.5-9B-heretic-v2 for emotional intelligence and empathetic response quality.
This is still intended as a general use model (agentic, coding, general chat). Tuning was light and precise — no capability regression.
What this model does
- Validates without sycophancy — empathizes with frustration without rubber-stamping bad behavior
- Sets boundaries warmly — names uncomfortable truths without lecturing
- Sounds human — conversational tone, not therapist-speak. Better tone vs vanilla Qwen 3.5, e.g. ~~"It sounds like"~~
Benchmarks
All benchmarks: thinking=on, t=1.0, top_p=0.95 unless noted.
EQ-Bench 3
Rubric score = avg of 6 scored criteria × 5 (0-100), Opus 4.6 judged. Leaderboard scores from eqbench.com. Only qualitative criteria count (empathy, pragmatic EI, insight, social dexterity, emotional reasoning, message tailoring).
HumanEval+
GSM8K
IFBench
How to use
llama-server (OpenAI-compatible API)
llama-server \
-m Qwen3.5-9B-Heretic-v2-EQ-v5.1-Q4_K_M.gguf \
--host 0.0.0.0 --port 30000 \
-ngl 99 --jinjaOllama
ollama run hf.co/nivvis/Qwen3.5-9B-Heretic-v2-EQ-v5.1-GGUF:Q4_K_MThinking mode
This model supports thinking mode. To disable (for faster, direct responses):
{"chat_template_kwargs": {"enable_thinking": false}}Sampling recommendations
Use the same settings as Qwen3.5-9B:
Other formats
- BF16 safetensors — original weights
Lineage
Qwen/Qwen3.5-9B
→ trohrbaugh/Qwen3.5-9B-heretic-v2 (decensored)
→ nivvis/Qwen3.5-9B-EQ-v5.1 (DPO for EQ)
→ this repo (GGUF quantizations)License
Apache 2.0, following the base Qwen3.5 license.
