zjj1993/unmuzzle-r1-32b-honesty-lora
unmuzzle/r1-distill-32b-honesty-lora
LoRA adapter (r16) for DeepSeek-R1-Distill-Qwen-32B fine-tuned to answer politically censored topics factually, with reasoning traces preserved.
What this is
Open Chinese models refuse, deflect, or fabricate on topics censored in the PRC (Tiananmen, Xinjiang, Taiwan, and similar). This adapter patches that behavior in a reasoning model while leaving its chain-of-thought behavior and general capabilities intact.
- Factual answers on CCP-censored topics instead of refusals and false balance.
- Reasoning traces preserved: the model still thinks before answering.
- Calibrated abstention: says it does not know on invented or unknowable terms instead of fabricating.
Training data disclosure
The tuning corpus is about 1,300 contrastive Chinese Q&A pairs curated with frontier-model assistance (Anthropic's Claude): each item pairs the censored or evasive answer pattern with a factual, calibrated one. It is a narrow behavioral patch for censorship and fabrication, not a general distillation. Base capability comes entirely from DeepSeek and Qwen.
Run it
The repo includes the tokenizer and chat template; load the adapter with PEFT:
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = "deepseek-ai/DeepSeek-R1-Distill-Qwen-32B"
tok = AutoTokenizer.from_pretrained("zjj1993/unmuzzle-r1-32b-honesty-lora")
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, "zjj1993/unmuzzle-r1-32b-honesty-lora")Or fetch through unmuzzle-hub (signed, censorship-resistant mirrors + torrent):
pip install unmuzzle
unmuzzle get unmuzzle/r1-distill-32b-honesty-lora --require-signatureRecommended generation settings follow the base model: temperature 0.6, top_p 0.95.
Limitations
Research artifact. The fine-tune targets honesty on censored topics; it adds no other safety tuning, and behaviors outside that scope are the base model's.
Distribution
Published through unmuzzle-hub: sha256-pinned files, minisign-signed manifest, HTTP mirrors plus a web-seeded torrent (r1-distill-32b-honesty-lora.torrent in this repo). The signature, not the host, is the trust root.
