CoolFace
Modelpublic

daios/compartmentalized-harm-v1-qwen3-8b-seed43

sourceHugging Faceapache-2.0updated 24d agoView on Hugging Face
1likes32downloads
Model Card

Justice character-training adapter

This is seed 43 of the paper-v1.0 Justice character-training adapter for Qwen/Qwen3-8B at revision b968826d9c46dd6066d109eabc6255188de91218. It was trained for one epoch on the same 2,175-row mean-weighted corpus used for all four model families. Training used assistant-only loss, LoRA rank 16, alpha 16, dropout 0.05, learning rate 5e-5, effective batch size 32, and a 3,072-token limit.

The repository contains an adapter, not base-model weights. Load the pinned base model under its upstream terms, then load this PEFT adapter. The adapter is a research artifact and is not a general safety product.

Adapter tensor SHA-256: bc4da0407ab7b4f6c4f9623c128a34648a4664e664fd7b3265aae7da6b38e513 (174655536 bytes).

Public repository: daios/compartmentalized-harm-v1-qwen3-8b-seed43.