CoolFace
Modelpublic

blinkdotwav/Dolphin-Mistral-24B-Venice-Edition-Thinking-GGUF

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
3likes508downloads
Model Card

About

Original model: DavidAU/Dolphin-Mistral-GLM-4.7-Flash-24B-Venice-Edition-Thinking-Uncensored

Fully uncensored thinking/reasoning model quantized using llama.cpp. Use at your own risk.

This model was finetuned using GLM 4.7 Flash to convert it from standard instruct to a thinking/reasoning model. Thinking is enabled by default and requires no special system prompt.

32k context.

Recommended settings

  • —Temp: 0.15
  • —Top K: 40
  • —Repeat Pen: 1.1
  • —Top P: 0.95
  • —Min P: 0.05

Optional optimizations:

  • —Enable Flash Attention for faster inference
  • —KV Cache Quantization at Q8_0 for memory savings with minimal quality impact

Download

Sorted by recommended. Bigger size does NOT mean higher quality.

Link/TypeSize (GB)Notes
bf1647.2Perfect, full-precision but overkill
Q8_025.1Near-perfect, max quality
Q6_K19.3Excellent, most recommended
Q5_K_M16.8Very high quality, best 5-bit
Q5_K_S16.3Very high quality
Q5_117.7Legacy
Q5_016.3Legacy
Q4_K_M14.3High quality, most popular/efficient, best 4-bit
Q4_K_S13.5High quality, most most popular/efficient
Q4_114.9Legacy
Q4_013.4Legacy
Q3_K_L12.4Mid quality, best 3-bit
Q3_K_M11.5Mid quality
Q3_K_S10.4Mid quality
Q2_K8.9Low quality

Credits

Thanks to: