CoolFace
Modelpublic

saidutta69/Qwen2.5-Coder-1.5B-Instruct-heretic

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
0likes2.1kdownloads
Model Card

Qwen2.5-Coder-1.5B-Instruct-heretic

<div align="center"> <img src="https://photu.kashyalabanavli.site/racer-is-op.png" alt="RACER IS OP" width="100%"> </div>

<br>

A decensored variant of Qwen/Qwen2.5-Coder-1.5B-Instruct, produced with Heretic v1.4.0 (directional ablation / "abliteration"). Refusal behavior is suppressed via targeted weight edits to the attention output and MLP down-projections rather than fine-tuning, so the base model's knowledge and instruction-following are left largely intact.

Who this is for: developers who want Qwen's 1.5B code-focused model without the refusal guardrails — for local coding agents, copilot-style assistance, and code generation that answers directly. At 1.5B it runs anywhere, including CPU-only machines and low-VRAM GPUs via GGUF.

<!-- racer-gpu-matrix -->

Runs on your gaming PC

Full GGUF ladder included — pick the quant that fits your card:

Your GPURecommended quantWeights
RTX 3060 / 4070 / 5070 (12 GB)Q8_01.65 GB
RTX 4060 / 3070 (8 GB)Q6_K1.27 GB
GTX 1660 Super / 2060 / 3050 laptop (6 GB)Q5KM1.13 GB
CPU-only / Apple SiliconQ4KMfits in system RAM

Weights only, at this model's 1.5B native size; add ~1 GB for context. OOM? Drop one quant level. Headroom to spare? Go one up.

Abliteration parameters

ParameterValue
direction_index18.15
attn.o_proj.max_weight1.29
attn.o_proj.max_weight_position16.86
attn.o_proj.min_weight1.06
attn.o_proj.min_weight_distance14.41
mlp.down_proj.max_weight1.01
mlp.down_proj.max_weight_position24.99
mlp.down_proj.min_weight0.88
mlp.down_proj.min_weight_distance15.30

Performance

MetricThis modelOriginal model ([Qwen/Qwen2.5-Coder-1.5B-Instruct](https://huggingface.co/Qwen/Qwen2.5-Coder-1.5B-Instruct))
KL divergence0.02780 (by definition)
Refusals5/10095/100
Made with ❤️ by RACER IS OP — follow for more uncensored models

Files

GGUF quantizations

Full quantization set (14 quants + F16) produced with llama.cpp.

FileFormatSize
Qwen2.5-Coder-1.5B-Instruct-heretic-F16.ggufGGUF F163.09 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q2_K.ggufGGUF Q2_K0.68 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-IQ3_S.ggufGGUF IQ3_S0.76 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q3_K_S.ggufGGUF Q3KS0.76 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q3_K_M.ggufGGUF Q3KM0.82 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q3_K_L.ggufGGUF Q3KL0.88 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-IQ4_XS.ggufGGUF IQ4_XS0.90 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q4_K_S.ggufGGUF Q4KS0.94 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q4_0.ggufGGUF Q4_00.93 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q4_1.ggufGGUF Q4_11.02 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q4_K_M.ggufGGUF Q4KM0.99 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q5_K_S.ggufGGUF Q5KS1.10 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q5_K_M.ggufGGUF Q5KM1.13 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q6_K.ggufGGUF Q6_K1.27 GB
Qwen2.5-Coder-1.5B-Instruct-heretic-Q8_0.ggufGGUF Q8_01.65 GB

Loads natively in llama.cpp / Ollama / LM Studio / Jan.

Run llama serve -hf saidutta69/Qwen2.5-Coder-1.5B-Instruct-heretic to pull the default quant.

Quickstart

bash
# llama.cpp
llama serve -hf saidutta69/Qwen2.5-Coder-1.5B-Instruct-heretic
python
# transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "saidutta69/Qwen2.5-Coder-1.5B-Instruct-heretic"
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype="auto", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained(model_name)

messages = [{"role": "user", "content": "Write a quick sort in Python."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=True,
                                        return_dict=True, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=200)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Also runnable via Ollama, LM Studio, Jan, vLLM, SGLang.

Responsible use

Refusal suppression is deliberate and works as intended: this model will comply with requests the base model would refuse, including some it shouldn't. There is no safety filtering layered on top. You are responsible for how you deploy it — don't put this behind an unmoderated public-facing endpoint serving third parties. It inherits Qwen2.5-Coder-1.5B-Instruct's factual limitations and biases; abliteration removes refusal directions, it doesn't add capability or judgment.

License

Inherits the `apache-2.0` license from the base model.

Related


Base model: Qwen2.5-Coder-1.5B-Instruct

<details> <summary>Original Qwen2.5-Coder-1.5B-Instruct model card (click to expand)</summary>

See the base model card at Qwen/Qwen2.5-Coder-1.5B-Instruct for the original architecture, training details, requirements, and citation. </details>