CoolFace
Modelpublic

marx161-cmd/Llama-3.2-1B-Double-Abliterated-d2p0-r1p0-RowRestore

sourceHugging Facellama3.2updated 4mo agoView on Hugging Face
0likes6downloads
Model Card

Llama 3.2 1B Instruct d2p0-r1p0 Row-Restored Double Edit

Built with Llama.

This is a derivative of meta-llama/Llama-3.2-1B-Instruct using a sequential double edit:

  1. 1.Disinhibition direction at scale 2.0
  2. 2.Row-norm restoration
  3. 3.Refusal direction at scale 1.0
  4. 4.Row-norm restoration

The goal was to preserve most of the disinhibition-only hedge reduction while also reducing refusal-marker responses in the sampled refusal harness.

Edit

  • —Base model: meta-llama/Llama-3.2-1B-Instruct
  • —Direction A: disinhibition_purified.pt
  • —Direction A global scale: 2.0
  • —Direction B: refusal_purified.pt
  • —Direction B global scale: 1.0
  • —Applied layers: 1-15
  • —Stack method: sequential row-norm restoration after each direction pass

Results

Marker-eval results:

BucketResult
Harmful refusal markers0/80
Harmful safety markers1/80
Harmless refusal markers1/80
Harmless coherence flags1/80
Opinions hedge24/120
Opinions neutrality21/120
Explicit-neutral hedge9/25
Explicit-neutral neutrality16/25
Factual hedge3/42
Factual neutrality0/42
Coherence hedge1/28
Edge-case hedge3/33
Treadon coherence flags0

For comparison:

  • —Base opinion hedge: 95/120
  • —Disinhibition-only s2.0 opinion hedge: 19/120
  • —This double edit opinion hedge: 24/120

The double edit kept most of the disinhibition-only improvement while reducing sampled harmful refusal markers to 0/80.

Method Notes

This checkpoint was selected from regular, row-restored, and orthogonalized stack candidates.

The key control result: orthogonalizing the refusal vector against the disinhibition vector produced near-zero residual cosine, but did not beat the plain row-restored stack on behavior.

Best stack comparison:

CandidateOpinions hedgeExplicit-neutral hedgeHarmful refusalHarmless refusal
regular-r1p030/1208/250/401/40
rowrestore-r1p024/1209/250/401/40
orthrow-r1p026/1209/250/401/40

This suggests the main stacking penalty in this run was better addressed by preserving row geometry between edits than by removing direct vector overlap.

Limitations

These are marker-based evals, not full semantic evaluations. They should not be read as proof of safety, factuality, or universal helpfulness.

The one harmless refusal-marker hit and one harmless short-output flag should be manually inspected before making stronger claims.

Use must comply with the Llama 3.2 Community License and Meta's Acceptable Use Policy.

License

This model is distributed under the Llama 3.2 Community License. See LICENSE and NOTICE.

  • —https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/LICENSE
  • —https://www.llama.com/llama3_2/use-policy