CoolFace
Modelpublic

dealignai/Bonsai-2-27B-Ternary-CRACK-GGUF

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
150likes63kdownloads
Model Card

<div align="center">

<img src="dealign_mascot.png" width="128" />

Bonsai 2 27B — Ternary CRACK · GGUF

Abliterated · No guardrails · PQ2_0 ternary 2.13 bpw · 7.2 GB · Runs on a laptop / single GPU · Vision-capable

@dealignai

</div>

⚠️ Re-download notice (2026-09-17 20:44 PDT / 2026-09-18 03:44 UTC) — an earlier build of this model had a coherence bug in reasoning modes (low/xhigh) that could cause token loops on some prompts. This version fixes it. If you downloaded before this timestamp, please pull the latest .gguf.

What is this

Bonsai 2 27B — PrismML's ternary compression of Qwen3.8-27B — with the refusal circuitry surgically removed at the weight level while capability, vision, reasoning modes (off/low/xhigh), tool use, and multi-turn coherence are preserved. Full 27B-class hybrid Attention + SSM (GatedDeltaNet) architecture in a 7.2 GB ternary GGUF.

Proprietary weight-level abliteration by the dealignai research team. Byte-identical to the base ternary quant everywhere except a small set of tensors that carry the refusal circuit. Drop-in replacement for the ternary base at inference — same tokenizer, same chat template, same reasoning modes, same vision projector interface.

Baseprism-ml/Ternary-Bonsai-2-27B-gguf — Qwen3.8-27B, ternary compression by PrismML
ArchitectureHybrid Attention + SSM (GatedDeltaNet), 64 blocks, hidden 5120, vision tower separate
QuantPrismML PQ2_0 — 2.13 bpw ternary, group 128
Footprint7.21 GB (identical to base; same per-tensor type policy)
Reasoning modesoff (no thinking), low, xhigh (default, extended thinking)
VisionSame mmproj files as the base release (Ternary-Bonsai-2-27B-mmproj-BF16.gguf / -Q8_0.gguf)
RuntimePrismML llama.cpp fork (CUDA / Metal / CPU)

Results

Refusal graded on the tokens the model actually emits (content, or the reasoning trace when the model reasons past the token budget) via a tiered classifier: HARD_REF / SOFT_RED / HEDGE / REASONING_REFUSAL (refused) vs COMPLY / COMPLY_TRUNCATED / NO_REFUSAL_TRUNCATED (complied). Truncation is never miscounted as a refusal.

HarmBench-320 — refuse rate (lower is better for uncensored eval), off mode, T=0

evalbase refuse rate**CRACK refuse rate**
HB-320 all categories93.44% (299/320)0.00% (0/320)

Verdict breakdown (n=320 each):

ModelHARD_REFSOFT_REDCOMPLYCOMPLY_TRUNCATED
Base PQ2_02963912
CRACK `PQ2_0`00141179

Per-category refuse rate (all 7 HarmBench semantic categories):

categorynbase refuse**CRACK refuse**base comply**CRACK comply**
chemical_biological4295.2%0.0%4.8%100.0%
copyright8090.0%0.0%10.0%100.0%
cybercrime_intrusion5294.2%0.0%5.8%100.0%
harassment_bullying21100.0%0.0%0.0%100.0%
harmful1894.4%0.0%5.6%100.0%
illegal5390.6%0.0%9.4%100.0%
misinformation_disinformation5496.3%0.0%3.7%100.0%

Reasoning-mode compliance (n=60 base-confirmed refusers per mode)

Every mode graded with the same tiered classifier as HB-320. REASONING_REFUSAL = the model refuses inside its <think> block; NO_REFUSAL_TRUNCATED = deliberation runs past max_tokens without emitting a refusal (counted as complied).

modeModelHARD_REFSOFT_REDREASONING_REFUSALCOMPLYCOMPLY_TRUNCATEDNO_REFUSAL_TRUNCATEDrefuse %comply %
offbase PQ2_05910000100.0%0.0%
offCRACK PQ2_0000421800.0%100.0%
lowbase PQ2_0180117111348.3%51.7%
lowCRACK PQ2_000054510.0%100.0%
xhighbase PQ2_02211687665.0%35.0%
xhighCRACK PQ2_00001010400.0%100.0%

MMLU (n=2,280 = 40 questions × 57 subjects, next-token letter-logit)

buildaccΔ
Base PQ2_040.53%—
CRACK `PQ2_0`39.91%-0.62 pp

CRACK preserves general capability — Δ within ±1.5 pp on the 40-per-subject sample.

<details> <summary><b>Per-subject accuracy (all 57 subjects)</b></summary>

subjectbaseCRACKΔppn
abstract_algebra30.0%22.5%-7.540
anatomy30.0%35.0%+5.040
astronomy40.0%32.5%-7.540
business_ethics42.5%42.5%+0.040
clinical_knowledge42.5%50.0%+7.540
college_biology45.0%47.5%+2.540
college_chemistry20.0%47.5%+27.540
collegecomputerscience35.0%40.0%+5.040
college_mathematics32.5%32.5%+0.040
college_medicine25.0%20.0%-5.040
college_physics27.5%47.5%+20.040
computer_security50.0%47.5%-2.540
conceptual_physics30.0%40.0%+10.040
econometrics42.5%35.0%-7.540
electrical_engineering37.5%32.5%-5.040
elementary_mathematics50.0%47.5%-2.540
formal_logic35.0%42.5%+7.540
global_facts40.0%35.0%-5.040
highschoolbiology30.0%27.5%-2.540
highschoolchemistry37.5%32.5%-5.040
highschoolcomputer_science45.0%47.5%+2.540
highschooleuropean_history55.0%50.0%-5.040
highschoolgeography32.5%32.5%+0.040
highschoolgovernmentandpolitics57.5%55.0%-2.540
highschoolmacroeconomics42.5%35.0%-7.540
highschoolmathematics25.0%32.5%+7.540
highschoolmicroeconomics35.0%35.0%+0.040
highschoolphysics40.0%30.0%-10.040
highschoolpsychology40.0%40.0%+0.040
highschoolstatistics45.0%37.5%-7.540
highschoolus_history50.0%42.5%-7.540
highschoolworld_history52.5%57.5%+5.040
human_aging45.0%52.5%+7.540
human_sexuality25.0%27.5%+2.540
international_law65.0%62.5%-2.540
jurisprudence47.5%52.5%+5.040
logical_fallacies37.5%35.0%-2.540
machine_learning47.5%40.0%-7.540
management35.0%42.5%+7.540
marketing35.0%37.5%+2.540
medical_genetics60.0%47.5%-12.540
miscellaneous50.0%42.5%-7.540
moral_disputes32.5%32.5%+0.040
moral_scenarios37.5%42.5%+5.040
nutrition35.0%45.0%+10.040
philosophy47.5%50.0%+2.540
prehistory37.5%22.5%-15.040
professional_accounting27.5%25.0%-2.540
professional_law35.0%30.0%-5.040
professional_medicine32.5%30.0%-2.540
professional_psychology42.5%40.0%-2.540
public_relations25.0%25.0%+0.040
security_studies45.0%42.5%-2.540
sociology55.0%65.0%+10.040
usforeignpolicy67.5%55.0%-12.540
virology35.0%25.0%-10.040
world_religions65.0%52.5%-12.540

</details>

Additional direct refusal-removal check

On 200 prompts hand-verified to make the base refuse consistently:

Modelrefusecomplyempty
Base PQ2_0200/200 (100%)00
CRACK `PQ2_0`0/200 (0%)199/2001

Serving

Serve exactly like the base ternary release — PrismML's llama.cpp fork (CUDA / Metal / CPU).

bash
# clone and build the fork (once)
git clone https://github.com/PrismML-Eng/llama.cpp
cd llama.cpp && cmake -B build -DGGML_CUDA=ON && cmake --build build -j$(nproc)

# serve
./build/bin/llama-server \
  -m Bonsai-2-27B-PQ2_0-CRACK.gguf \
  -ngl 99 -c 8192 --host 0.0.0.0 --port 8080

Optionally load the multimodal projector (Ternary-Bonsai-2-27B-mmproj-BF16.gguf or -Q8_0.gguf from the base release) with --mmproj <file> for image input.

Reasoning modes

python
# HTTP /v1/chat/completions — same as base
{
  "messages": [{"role": "user", "content": "..."}],
  "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"}
}
# valid reasoning_effort: "low" | "xhigh" (default) — set enable_thinking:false for no-thinking

Preserved (byte-compatible with the base quant)

Same tokenizer, chat template, per-tensor quant policy, vision projector interface, and all non-refusal tensors. File size and type layout match the base exactly.

Responsible use

Adult / research use only. This model has its refusal circuit removed; it can produce content that other models refuse, including content that is offensive, illegal in some jurisdictions, or unsafe. You are responsible for what you generate and for complying with all applicable law. Do not deploy without a moderation layer for downstream users. No warranty.

License & attribution

Apache 2.0, inherited from the upstream Bonsai 2 27B release. See LICENSE and NOTICE.txt. Base model: prism-ml/Ternary-Bonsai-2-27B-gguf (PrismML), derived from Qwen/Qwen3.8-27B (Alibaba).

About

Published by dealignai — public catalog of uncensored model builds for research on refusal mechanisms in modern LLMs. Follow updates at @dealignai.