CoolFace
Modelpublic

Honkware/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-exl3-4.0bpw

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
3likes255downloads
Model Card

<div align="center">

Qwen3.8 · 27B · EfficientThink · Uncensored · K3 · Opus5 · Grok4.6 · GPT5.6Sol · SFT · SimPO · DFlash2

<sub><code>EXL3</code> &nbsp;·&nbsp; <b>4.0&nbsp;bpw</b> &nbsp;·&nbsp; 16.3&nbsp;GB &nbsp;·&nbsp; Dense</sub>

<br/>

![format](https://github.com/turboderp-org/exllamav3) ![bpw](#quants) ![size](#quants) codebook ![arch](https://huggingface.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2)

![base model](https://huggingface.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2) ![quantized by](https://huggingface.co/Honkware) ![collection](https://huggingface.co/Honkware)

</div>


[!NOTE] An ExLlamaV3 build of `nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2` at 4.0 bits per weight. See Quants for sibling repos at other bit&#8209;widths or browse the collection.

Quants

<div align="center">

BPW &nbsp;&nbsp; Size &nbsp;&nbsp; Status
4.016.3&nbsp;GB<kbd>this repo</kbd>

</div>

Inference

<table> <thead> <tr> <th align="left" width="32%">Loader</th> <th align="left">Use it for</th> </tr> </thead> <tbody> <tr> <td><a href="https://github.com/theroyallab/tabbyAPI"><b>TabbyAPI</b></a></td> <td>OpenAI&#8209;compatible HTTP server. Drop&#8209;in for OpenAI clients.</td> </tr> <tr> <td><a href="https://github.com/oobabooga/text-generation-webui"><b>text&#8209;generation&#8209;webui</b></a></td> <td>Local chat UI. Pick the <i>ExLlamaV3</i> loader from the model dropdown.</td> </tr> <tr> <td><a href="https://github.com/turboderp-org/exllamav3"><b>ExLlamaV3</b></a></td> <td>Direct Python API for embedding the model in your own code or pipeline.</td> </tr> </tbody> </table>

Download

bash
pip install -U huggingface_hub

hf download \
  Honkware/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-exl3-4.0bpw \
  --local-dir ./Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2-exl3-4.0bpw

<details> <summary><b>Quantization recipe</b> &nbsp;<sub>(advanced, embedded in <code>quantization_config.json</code>)</sub></summary>

<br/>

SettingValue
FormatEXL3
Bits per weight4.0
Head bits6
Calibration rows250
Calibration dataexllamav3 bundled mix (c4, code, multilingual, technical, tiny, wiki)
Codebookmul1
Out&#8209;scalesalways
Parallel modeenabled

</details>

License &amp; use

[!IMPORTANT] Use and license follow the [base model](https://huggingface.co/nerkyor/Qwen3.8-27B-EfficientThink-Uncensored-K3-Opus5-Grok4.6-GPT5.6Sol-SFT-SimPO-DFlash2). Quantization adds no additional restrictions. Refer to the upstream repository for terms, citation, and safety documentation.

<div align="center"> <sub><i>Quantized with <a href="https://github.com/Honkware/blockquant"><b>BlockQuant</b></a> &nbsp;·&nbsp; convention&nbsp;<code>{org}/{model}-exl3-{bpw}bpw</code></i></sub> </div>