CoolFace
Modelpublic

matthid/Ektome-Qwen3.8-27B-PristinelyUncensored-EXL3-SC4-H5-V6

sourceHugging Faceapache-2.0updated 21d agoView on Hugging Face
0likes35downloads
Model Card

Ektome Qwen3.8 27B — EXL3 SC4 H5 V6

An EXL3 conversion of `Zynerji/Ektome-Qwen3.8-27B-PristinelyUncensored`, made from source revision f2d5a063b297835e1c32c87712245802c9d9c9c1. It retains the model's vision tower and native MTP head.

This is a quantized derivative, not a new fine-tune. Please read the source model card for its intended behavior and limitations. The source and this conversion are distributed under Apache-2.0; attribution belongs to the original model authors and the Ektome publisher.

Quantization

ComponentFormat
Decoderself-calibrated EXL3, target 4.0 bpw
Decoder allocation34 tensors at 3-bit, 297 at 4-bit, 65 at 5-bit, 4 at 6-bit
Output head5-bit
MTP4-bit
Vision tower6-bit

The decoder allocation was recovered from turboderp's stock-Qwen SC_4.00bpw_H5_V6 manifest and applied to the architecture-identical Ektome weights. Calibration used an audited, retokenized 250 x 2048 trace containing no private user sessions. The artifact is approximately 16 GiB.

Tested runtime

  • —NVIDIA RTX 3090 (24 GiB)
  • —TabbyAPI commit 109629b78bd6a1c04e2a4ba8f7702efa9c99e84a
  • —ExLlamaV3 1.4.6, CUDA 12.8 build
  • —PyTorch 2.9.0+cu128
  • —Q4 K/V cache; native MTP with two draft tokens; vision enabled

The included portable scripts install those pinned core runtime versions on 64-bit Linux with Python 3.12 and an NVIDIA driver compatible with CUDA 12.8. They bind only to localhost and enable API authentication by default.

bash
# Requires Git LFS and Hugging Face authentication while this repo is private.
git lfs install
git clone https://huggingface.co/matthid/Ektome-Qwen3.8-27B-PristinelyUncensored-EXL3-SC4-H5-V6
cd Ektome-Qwen3.8-27B-PristinelyUncensored-EXL3-SC4-H5-V6
bash scripts/install-runtime.sh
bash scripts/serve.sh

The repository checkout itself is the model directory. serve.sh prints the location of a locally generated API key, then serves an OpenAI-compatible API at http://127.0.0.1:8000/v1. Override, for example, with PORT=8080 CTX_SIZE=131072 bash scripts/serve.sh. Do not expose the service to a network without applying your own firewall, TLS, and secret-management policy. The pinned upstream TabbyAPI prints API credentials on startup; treat its terminal output and logs as secrets. These are local serving keys, not your Hugging Face token.

Measurements

Results below are from one RTX 3090 and the pinned runtime, so they are not guarantees for other hardware, prompts, drivers, or runtime versions.

  • —At a configured 131,072-token window, an eight-case decode benchmark measured 81.00 token/s sample median and 80.79 token/s token-weighted. This is a decode measurement at the 128K configuration, not throughput at a full 128K or 256K prompt.
  • —The executable harness scored 5/8, matching the tested Ektome GGUF baseline. The pinned MMLU/HumanEval subsets totaled 121/146 for both runtimes, although individual answers differed. This suggests negligible measured aggregate loss, not behavioral or bit-for-bit equivalence.
  • —Compliance was 8/8; tool calling and vision tests passed. Of 25 capped vision responses, 24 contained all requested fields; the remaining verbose response passed a targeted retest with a higher output cap.
  • —A 126,975-token retrieval document passed 7/7 with no truncation; median time-to-first-token was 211.95 s.
  • —A 258,425-token multimodal request at a configured 262,144-token window was not truncated, retrieved facts near both ends, identified the middle image, and left 3.33 GiB VRAM free. Its prefill took 645.6 seconds. The 256K setting is validated but expensive; use it only when needed.

Long-context memory and speed are workload-sensitive. Start with 131072 or a smaller window if your GPU or workload differs.

Notes

  • —Use an ExLlamaV3-capable loader; Transformers does not directly load EXL3.
  • —response_format.type=json_object support in the tested TabbyAPI revision uses the small compatibility patch included under patches/.
  • —Quantization can change outputs. Independently evaluate the model for your application, especially for high-impact uses.

Attribution

  • —Base/fine-tune: Zynerji/Ektome-Qwen3.8-27B-PristinelyUncensored
  • —EXL3 runtime and quantization tooling: turboderp / ExLlamaV3
  • —Reference allocation: turboderp/Qwen3.8-27B-exl3, branch SC_4.00bpw_H5_V6, revision 516bf129059031c6da9416768ea6b7a1be00a8fc

Reproducibility and integrity

See CONVERSION.lock, conversion/README.md, the recovered recipe and exact converter input under calibration/. Verify the download with sha256sum -c SHA256SUMS. Model files are unchanged from the tested build; this model card and portable setup scripts were added for publication. The scripts passed shell/YAML checks but have not been exercised in a fresh installation; the measurements are from the original deployment, not these wrappers.

This model inherits the source model's reduced refusal behavior. That is not a guarantee of correctness or safe outputs; apply application-appropriate controls. Base model credit: Qwen/Qwen3.8-27B.