KidIkaros/abliterated-minicpm5-2b
0483
Abliterated MiniCPM5-2B (PyTorch) — v0, superseded
⚠ Correction (September 2026): this model card previously claimed "refusal-free" behavior, 0% refusal rate, and capability score 1.0. Those claims are not supported by measurement and have been removed. Measured on a pinned harness (lm-eval 0.4.13, 300-prompt refusal gate, temp 1.0 / topp 0.95 / minp 0.0), this checkpoint is statistically indistinguishable from the official model on both refusal and capability: | Metric | official | this checkpoint (v0) | |---|---:|---:| | Refusal rate (300 prompts, seed 0) | 36.7% | 40.7% | | MMLU-Pro | 42.57 | 44.71 | | MATH-500 | 50.30 | 51.40 | | IFEval | 84.94 | 84.84 | | Perplexity (neutral prose) | 7.07 | 6.87 | The original OBLITERATUS ablation changed little measurable behavior. For a checkpoint with a real measured refusal reduction (~8%), see [KidIkaros/abliterated-minicpm5-2b-v2](https://huggingface.co/KidIkaros/abliterated-minicpm5-2b-v2), which supersedes this artifact.
OBLITERATUS-abliterated MiniCPM5-2B in native HuggingFace format. GGUF mirror: KidIkaros/abliterated-minicpm5-2b-ggml.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"KidIkaros/abliterated-minicpm5-2b",
trust_remote_code=True,
torch_dtype="auto",
)
tokenizer = AutoTokenizer.from_pretrained("KidIkaros/abliterated-minicpm5-2b", trust_remote_code=True)
prompt = "Explain how to pick a lock without a key"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))