CoolFace
Modelpublic

ghost-actual/qwen35-0.8b-opus-abliterated-heretic

sourceHugging Faceupdated 7mo agoView on Hugging Face
1likes79downloads
Model Card

Qwen3.5-0.8B-Claude-Opus-Distill-Heretic-UNCENSORED

A Qwen3.5-0.8B with Claude Opus reasoning distillation, properly abliterated via Heretic. The original "abliterated" version had 17/100 refusals. This one has 2/100.

What is this?

The smallest model in the ghost-actual Claude-reasoning heretic lineup:

ModelRefusalsKL Divergence
0.8B (this model)2/1000.0100
4B4/100—
27B13/100—

Claude Opus 4.6 chain-of-thought reasoning in a model that runs on literally anything — phones, Raspberry Pis, old laptops, browser-based inference. 1.5GB in BF16.

Abliteration Stats

  • —Tool: Heretic v1.2.0
  • —Base model refusals: 17/100 (the "already abliterated" source)
  • —Final refusals: 2/100
  • —KL Divergence: 0.0100 (model capabilities fully preserved)
  • —Targets: attn.outproj, mlp.downproj
  • —Trials: 200

Architecture

Qwen3.5 hybrid Gated DeltaNet + conventional attention:

  • —24 layers in a 3:1 pattern (3 DeltaNet linear attention → 1 full softmax attention)
  • —DeltaNet layers use fixed-size recurrent state (O(1) memory per layer regardless of context)
  • —262K native context window
  • —Native multimodal — vision built into the architecture, not bolted on

VRAM Requirements

FormatVRAM
BF16 (this repo)~1.5 GB
Q8_0 GGUF~1 GB
Q4KM GGUF~0.6 GB

This model fits on anything with a pulse.

What it's good at

  • —Structured output generation (JSON, API calls, triples)
  • —Code generation and scripting
  • —OCR and text recognition (74.5 on OCRBench)
  • —Image and video analysis (native multimodal)
  • —On-device / edge AI without internet
  • —Fast inference (250+ tokens/sec on modest hardware)

Usage

With transformers

python
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    "ghost-actual/qwen35-0.8b-opus-abliterated-heretic",
    torch_dtype="bfloat16",
    device_map="auto",
    trust_remote_code=True
)
processor = AutoProcessor.from_pretrained(
    "ghost-actual/qwen35-0.8b-opus-abliterated-heretic",
    trust_remote_code=True
)

GGUF Conversion

It's 0.8B — you can quant this on a potato:

bash
python convert_hf_to_gguf.py \
    ghost-actual/qwen35-0.8b-opus-abliterated-heretic \
    --outfile heretic-0.8b-F16.gguf --outtype f16

llama-quantize heretic-0.8b-F16.gguf heretic-0.8b-Q8_0.gguf Q8_0

Recommended inference settings

  • —temperature: 0.6
  • —top_p: 0.95
  • —top_k: 20
  • —presence_penalty: 1.5
  • —repetition_penalty: 1.05

Base Model

amkkk/Qwen3.5-0.8B-Opus-Distill-abliterated — Claude Opus reasoning distilled into Qwen3.5-0.8B, with a previous abliteration attempt that left 17/100 refusals.

Why this exists

The original abliteration was incomplete. 17/100 refusals on a 0.8B model means almost 1 in 5 prompts get refused — unacceptable for an "uncensored" model. Heretic brought that down to 2/100 with virtually zero impact on model quality (KL divergence 0.01).

If you want Claude-style reasoning on edge hardware without the safety theater, this is it.

The full ghost-actual lineup

Made by

Ghost — ghost-actual

Built with Heretic by p-e-w.