ghost-actual/qwen35-0.8b-opus-abliterated-heretic
Qwen3.5-0.8B-Claude-Opus-Distill-Heretic-UNCENSORED
A Qwen3.5-0.8B with Claude Opus reasoning distillation, properly abliterated via Heretic. The original "abliterated" version had 17/100 refusals. This one has 2/100.
What is this?
The smallest model in the ghost-actual Claude-reasoning heretic lineup:
Claude Opus 4.6 chain-of-thought reasoning in a model that runs on literally anything — phones, Raspberry Pis, old laptops, browser-based inference. 1.5GB in BF16.
Abliteration Stats
- Tool: Heretic v1.2.0
- Base model refusals: 17/100 (the "already abliterated" source)
- Final refusals: 2/100
- KL Divergence: 0.0100 (model capabilities fully preserved)
- Targets: attn.outproj, mlp.downproj
- Trials: 200
Architecture
Qwen3.5 hybrid Gated DeltaNet + conventional attention:
- 24 layers in a 3:1 pattern (3 DeltaNet linear attention → 1 full softmax attention)
- DeltaNet layers use fixed-size recurrent state (O(1) memory per layer regardless of context)
- 262K native context window
- Native multimodal — vision built into the architecture, not bolted on
VRAM Requirements
This model fits on anything with a pulse.
What it's good at
- Structured output generation (JSON, API calls, triples)
- Code generation and scripting
- OCR and text recognition (74.5 on OCRBench)
- Image and video analysis (native multimodal)
- On-device / edge AI without internet
- Fast inference (250+ tokens/sec on modest hardware)
Usage
With transformers
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model = Qwen3_5ForConditionalGeneration.from_pretrained(
"ghost-actual/qwen35-0.8b-opus-abliterated-heretic",
torch_dtype="bfloat16",
device_map="auto",
trust_remote_code=True
)
processor = AutoProcessor.from_pretrained(
"ghost-actual/qwen35-0.8b-opus-abliterated-heretic",
trust_remote_code=True
)GGUF Conversion
It's 0.8B — you can quant this on a potato:
python convert_hf_to_gguf.py \
ghost-actual/qwen35-0.8b-opus-abliterated-heretic \
--outfile heretic-0.8b-F16.gguf --outtype f16
llama-quantize heretic-0.8b-F16.gguf heretic-0.8b-Q8_0.gguf Q8_0Recommended inference settings
- temperature: 0.6
- top_p: 0.95
- top_k: 20
- presence_penalty: 1.5
- repetition_penalty: 1.05
Base Model
amkkk/Qwen3.5-0.8B-Opus-Distill-abliterated — Claude Opus reasoning distilled into Qwen3.5-0.8B, with a previous abliteration attempt that left 17/100 refusals.
Why this exists
The original abliteration was incomplete. 17/100 refusals on a 0.8B model means almost 1 in 5 prompts get refused — unacceptable for an "uncensored" model. Heretic brought that down to 2/100 with virtually zero impact on model quality (KL divergence 0.01).
If you want Claude-style reasoning on edge hardware without the safety theater, this is it.
The full ghost-actual lineup
- 0.8B (this model) — runs on anything
- 4B — lightweight desktop
- 27B — full power
- 27B GGUF — quantized for single-GPU setups
Made by
Ghost — ghost-actual
Built with Heretic by p-e-w.
