davetha/Ternary-Bonsai-2-27B-Abliterated-PQ2_0-GGUF
Ternary-Bonsai-2-27B Abliterated — PQ2_0
A format repack, not a new ablation. Same weights as BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-PTQ1_0-GGUF, in PQ20 instead of PTQ10, because PTQ1_0 has GPU kernels for NVIDIA Ampere only.
On an AMD MI210 (gfx90a / CDNA2) that is a 2.2x difference:
Single-stream, -ngl 99 -c 8192, one MI210, PrismML llama.cpp fork.
Why PTQ1_0 is slow outside Ampere
MMQ tile configs in the PrismML fork exist per architecture. PTQ1_0 has entries for Ampere and nowhere else:
mmq-config-ampere.cuh PTQ1_0=11 PQ2_0=16 Q2_0=16
mmq-config-cdna.cuh PTQ1_0=0 PQ2_0=7 Q2_0=7
mmq-config-rdna2/3/4 PTQ1_0=0 PQ2_0=12 Q2_0=12
mmq-config-pascal.cuh PTQ1_0=0 PQ2_0=11 Q2_0=11PTQ10 is additionally `#if !defined(GGMLUSEHIP)`-gated in 17 places, and its eligibility test is `turingmmaavailable(cc)`, hard-wired to NVIDIA. So on CDNA it falls back to fp16 dequantise + BLAS. PQ20 has CDNA kernels and takes the fast path.
How it was made
The ablated weights ship only as PTQ1_0, so the edit was transplanted rather than re-derived:
- Dequantised the abliterated and stock PTQ10 checkpoints and diffed them. **Exactly 98 tensors differ** — `ffndown
,ssmout`, `attnoutput`, blocks 15-63 — matching the source release's stated scope. - Added that delta to the F16 checkpoint:
W_new = W_stock_f16 + (W_ref - W_stock). - Quantised F16 -> PQ2_0.
Do not try to reproduce this by projecting a refusal direction out. The source card explains why, and it is worth repeating: on a ternary lattice, W <- W - λ·r·(rᵀW) followed by a repack changes 0 of 89,128,960 digits — the correction is ~1.4% of a weight's magnitude while the lattice step is 100%, so every weight rounds straight back. A projected-then-repacked model passes a file-size check and a quantiser reproducibility check while containing no ablation at all.
I made exactly that mistake on the first attempt here, inferring a projection from the phrase "rank-1". Weight-space measurement caught it: the reference's projection onto the recovered direction is ~3x higher than stock (0.047 vs 0.015), not zero — the real edit flips lattice digits, it does not project.
Verification
Weight-exact against the source release. Dequantised and compared per tensor:
Byte-level diff against stock run through the identical quantiser: 98 tensors differ (0.56-0.65% of bytes), 752 byte-identical. Quantisation is deterministic here, so that isolates the edit — it survived requantisation rather than being rounded away.
Because the weights are numerically identical to the source release, that release's evaluations apply — refusal 83.0% -> 1.0% on SimpleSafetyTests, XSTest-safe over-refusal 1.6% -> 0.0%, HumanEval pass@1 0.811 -> 0.805 (McNemar p = 1.000). Those numbers are BoldingBuilds', measured on the PTQ1_0 weights; I did not re-run them. What I verified is that these weights are the same ones, plus a capability spot-check (4/4 factual/arithmetic, coherent long-form).
PQ20 here is 2.45 bpw vs the stock PQ20 release's 2.13, because some tensors take a higher-precision fallback through this path.
Requirements
Needs the PrismML llama.cpp fork (PrismML-Eng/llama.cpp, branch prism). Stock llama.cpp cannot load PQ2_0. For CDNA2:
cmake -S . -B build -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx90a \
-DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
cmake --build build -jIt is a reasoning model: it emits reasoning_content before the answer, so a small max_tokens truncates inside the reasoning and returns empty content. Give it room.
Attribution
Ablation by BoldingBuilds (Apache 2.0) — this repo only changes the quantisation format. Created using Bonsai by Prism ML. Built from Qwen3.8-27B, Copyright 2026 Alibaba Cloud (Apache 2.0). See LICENSE and NOTICE.txt.
