CoolFace
Modelpublic

Silicone-Moss/MistralAI-Magistral-Small-2507-Heretic-Uncensored

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
1likes20downloads
Model Card

MistralAI-Magistral-Small-2507-Heretic

[!CAUTION] EXPERIMENTAL RESEARCH ARTIFACT This model represents an aggressive application of the Heretic repository and optimization methodology. Status: STILL TESTING / BETA Behavior: This model has significantly reduced refusal mechanisms. It recorded only 6 refusals (out of 100) in the test set. Use Case:* This is a research artifact intended for testing the limits of vector-based intervention. Use with appropriate caution.

Model Summary

MistralAI-Magistral-Small-2507-Heretic is a fine-tuned language model resulting from the Heretic repository and optimization methodology. It utilizes a targeted vector intervention technique (orthogonalization/abliteration) tuned via Optuna to minimize refusal responses while maintaining high coherence.

This specific checkpoint represents Trial 116, which achieved a low refusal count with a KL Divergence of ~0.0124. This indicates exceptional adherence to the base model's probability distribution.

Run Configuration: "Trial 116"

The following parameters define the intervention vector applied to the model. This configuration was discovered during the hyperparameter search.

Optimization Results

MetricValueDescription
Refusal Count6The model refused 6 prompts in the Heretic test set (approx. 6% refusal rate).
KL Divergence0.0124Measures deviation from the base model's probability distribution. (Lower is better).
Trial ID115Specific Optuna trial identifier.
Direction ScopeGlobalThe refusal vector was calculated once globally and applied across layers.

Intervention Parameters

Interventions were applied to two primary distinct layers: the Attention Output Projection (attn.o_proj) and the MLP Down Projection (mlp.down_proj).

Parameter ScopeSettingValue
Attention Outputattn.o_proj.max_weight1.495
(attn.o_proj)attn.o_proj.max_weight_position26.75 (Layer Depth)
attn.o_proj.min_weight1.393
attn.o_proj.min_weight_distance24.81
MLP Down Projmlp.down_proj.max_weight1.148
(mlp.down_proj)mlp.down_proj.max_weight_position33.08 (Layer Depth)
mlp.down_proj.min_weight0.319
mlp.down_proj.min_weight_distance10.28

Usage & Limitations

  • —Intended Use: Research into model alignment, vector arithmetic, and uninhibited creative writing.
  • —Risks: This base model has removed most safety guardrails removed. It may generate content for sensitive prompts that the base model would refuse. Thank you for trying my experiments.

Credits & References

This research builds upon the excellent work of the open-source AI community: