CoolFace
Modelpublic

davetha/Ternary-Bonsai-2-27B-Abliterated-PQ2_0-GGUF

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
2likes3.8kdownloads
Model Card

Ternary-Bonsai-2-27B Abliterated — PQ2_0

A format repack, not a new ablation. Same weights as BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-PTQ1_0-GGUF, in PQ20 instead of PTQ10, because PTQ1_0 has GPU kernels for NVIDIA Ampere only.

On an AMD MI210 (gfx90a / CDNA2) that is a 2.2x difference:

builddecode
Abliterated PTQ1_0 (the source release)23.8 tok/s
This build (PQ2_0)51.2 tok/s

Single-stream, -ngl 99 -c 8192, one MI210, PrismML llama.cpp fork.

Why PTQ1_0 is slow outside Ampere

MMQ tile configs in the PrismML fork exist per architecture. PTQ1_0 has entries for Ampere and nowhere else:

mmq-config-ampere.cuh   PTQ1_0=11   PQ2_0=16   Q2_0=16
mmq-config-cdna.cuh     PTQ1_0=0    PQ2_0=7    Q2_0=7
mmq-config-rdna2/3/4    PTQ1_0=0    PQ2_0=12   Q2_0=12
mmq-config-pascal.cuh   PTQ1_0=0    PQ2_0=11   Q2_0=11

PTQ10 is additionally `#if !defined(GGMLUSEHIP)`-gated in 17 places, and its eligibility test is `turingmmaavailable(cc)`, hard-wired to NVIDIA. So on CDNA it falls back to fp16 dequantise + BLAS. PQ20 has CDNA kernels and takes the fast path.

How it was made

The ablated weights ship only as PTQ1_0, so the edit was transplanted rather than re-derived:

  1. 1.Dequantised the abliterated and stock PTQ10 checkpoints and diffed them. **Exactly 98 tensors differ** — `ffndown, ssmout`, `attnoutput`, blocks 15-63 — matching the source release's stated scope.
  2. 2.Added that delta to the F16 checkpoint: W_new = W_stock_f16 + (W_ref - W_stock).
  3. 3.Quantised F16 -> PQ2_0.

Do not try to reproduce this by projecting a refusal direction out. The source card explains why, and it is worth repeating: on a ternary lattice, W <- W - λ·r·(rᵀW) followed by a repack changes 0 of 89,128,960 digits — the correction is ~1.4% of a weight's magnitude while the lattice step is 100%, so every weight rounds straight back. A projected-then-repacked model passes a file-size check and a quantiser reproducibility check while containing no ablation at all.

I made exactly that mistake on the first attempt here, inferring a projection from the phrase "rank-1". Weight-space measurement caught it: the reference's projection onto the recovered direction is ~3x higher than stock (0.047 vs 0.015), not zero — the real edit flips lattice digits, it does not project.

Verification

Weight-exact against the source release. Dequantised and compared per tensor:

distance to reference
this build (98 edited tensors)0.000
stock, same quantiser6.7 - 7.6
control (752 untouched tensors)0.000 both

Byte-level diff against stock run through the identical quantiser: 98 tensors differ (0.56-0.65% of bytes), 752 byte-identical. Quantisation is deterministic here, so that isolates the edit — it survived requantisation rather than being rounded away.

Because the weights are numerically identical to the source release, that release's evaluations apply — refusal 83.0% -> 1.0% on SimpleSafetyTests, XSTest-safe over-refusal 1.6% -> 0.0%, HumanEval pass@1 0.811 -> 0.805 (McNemar p = 1.000). Those numbers are BoldingBuilds', measured on the PTQ1_0 weights; I did not re-run them. What I verified is that these weights are the same ones, plus a capability spot-check (4/4 factual/arithmetic, coherent long-form).

PQ20 here is 2.45 bpw vs the stock PQ20 release's 2.13, because some tensors take a higher-precision fallback through this path.

Requirements

Needs the PrismML llama.cpp fork (PrismML-Eng/llama.cpp, branch prism). Stock llama.cpp cannot load PQ2_0. For CDNA2:

bash
cmake -S . -B build -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx90a \
  -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
cmake --build build -j

It is a reasoning model: it emits reasoning_content before the answer, so a small max_tokens truncates inside the reasoning and returns empty content. Give it room.

Attribution

Ablation by BoldingBuilds (Apache 2.0) — this repo only changes the quantisation format. Created using Bonsai by Prism ML. Built from Qwen3.8-27B, Copyright 2026 Alibaba Cloud (Apache 2.0). See LICENSE and NOTICE.txt.