CoolFace
Modelpublic

nivvis/Qwen3.5-9B-EQ-v5.1-GGUF

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
3likes220downloads
Model Card

Qwen3.5-9B-Heretic-v2-EQ-v5.1-GGUF

GGUF quantizations of nivvis/Qwen3.5-9B-EQ-v5.1 for llama.cpp, Ollama, and LM Studio.

Example outputs: EQ 9B v5.1 vs Vanilla Qwen3.5-9B — same prompt, same settings.

Available quantizations

QuantSizeNotes
F16~17.9 GBFull precision, lossless conversion
Q4KM~5.6 GBBest 4-bit balance, recommended for most users

Qwen3.5-9B-EQ-v5.1

A DPO fine-tune of trohrbaugh/Qwen3.5-9B-heretic-v2 for emotional intelligence and empathetic response quality.

This is still intended as a general use model (agentic, coding, general chat). Tuning was light and precise — no capability regression.

What this model does

  • —Validates without sycophancy — empathizes with frustration without rubber-stamping bad behavior
  • —Sets boundaries warmly — names uncomfortable truths without lecturing
  • —Sounds human — conversational tone, not therapist-speak. Better tone vs vanilla Qwen 3.5, e.g. ~~"It sounds like"~~

Benchmarks

All benchmarks: thinking=on, t=1.0, top_p=0.95 unless noted.

EQ-Bench 3

Rubric score = avg of 6 scored criteria × 5 (0-100), Opus 4.6 judged. Leaderboard scores from eqbench.com. Only qualitative criteria count (empathy, pragmatic EI, insight, social dexterity, emotional reasoning, message tailoring).

#ModelScore
1gpt-5.480.9
2claude-sonnet-4-679.9
3claude-opus-4-678.6
4gemma-4-31B-it72.8
5Qwen3.5-397B70.1
6Qwen3.5-35b-EQ-v5.068.2
7Qwen3.5-35b-EQ-v5.166.8
8Qwen3.5-9b-EQ-v5.165.4
9Qwen3-235B61.4
10gpt-4.5-preview59.7
11Qwen3.5-9b (vanilla)59.1
12o4-mini58.1
13DeepSeek-V3-032457.9
14Qwen3.5-35b (vanilla)55.4
15gpt-oss-120b51.4

HumanEval+

ModelHumanEval baseHumanEval+
EQ 9B v5.192.7%86.0%
Vanilla Qwen3.5-9B93.9%87.2%

GSM8K

MetricEQ 9B v5.1Vanilla Qwen3.5-9B
Accuracy84.9%82.8%

IFBench

MetricEQ 9B v5.1Vanilla Qwen3.5-9B
Loose (leaderboard)67.3%65.0%
Strict56.8%56.8%

How to use

llama-server (OpenAI-compatible API)

bash
llama-server \
  -m Qwen3.5-9B-Heretic-v2-EQ-v5.1-Q4_K_M.gguf \
  --host 0.0.0.0 --port 30000 \
  -ngl 99 --jinja

Ollama

ollama run hf.co/nivvis/Qwen3.5-9B-Heretic-v2-EQ-v5.1-GGUF:Q4_K_M

Thinking mode

This model supports thinking mode. To disable (for faster, direct responses):

json
{"chat_template_kwargs": {"enable_thinking": false}}

Sampling recommendations

Use the same settings as Qwen3.5-9B:

Modetemptop_ptop_kpresence_penalty
Thinking (general)1.00.95201.5
Thinking (coding)0.60.95200.0
Non-thinking (general)0.70.8201.5
Non-thinking (reasoning)1.00.95201.5

Other formats

Lineage

Qwen/Qwen3.5-9B
  → trohrbaugh/Qwen3.5-9B-heretic-v2 (decensored)
    → nivvis/Qwen3.5-9B-EQ-v5.1 (DPO for EQ)
      → this repo (GGUF quantizations)

License

Apache 2.0, following the base Qwen3.5 license.