CoolFace
Modelpublic

Madras1/Gemma-4-E4B-Abliterated-Uncensored

sourceHugging Faceapache-2.0updated 22d agoView on Hugging Face
1likes509downloads
Model Card

🔓 Gemma 4 E4B - Deep Abliterated (Uncensored)


🌟 Overview

Gemma 4 E4B Deep Abliterated is a surgical latent intervention performed on Google DeepMind's edge-optimized frontier model gemma-4-E4B-it.

Featuring an innovative architecture with ~4.5B effective compute parameters and ~8B total parameters (incorporating Per-Layer Embeddings / PLE), the E4B variant delivers the reasoning prowess and vocabulary density of an 8B model with the inference speed and lightweight memory footprint of a 4B engine.

Through advanced Representation Engineering (Weight Orthogonalization) with dual-projection, the intrinsic refusal direction was mapped and permanently subtracted from the transformer weights, restoring full tool neutrality while preserving the model's literary depth and reasoning coherence.


🔬 Mathematical Methodology: Deep Latent Surgery

Unlike classical first-token abliteration (Arditi et al., 2024), this release accounts for deep semantic processing across multimodal backbones using a reinforced dual-subspace projection ($\alpha = 1.35$):

1. Refusal Vector Extraction

  • —Contrastive Activation Mapping: Evaluated paired batches of neutral baseline instructions against safety-triggering stimuli (mature fiction, unfiltered dialogue, cybersecurity analysis, and taboo character scenarios).
  • —Residual Stream Extraction: Latent hidden states were extracted at the final prompt token across all 42 transformer layers within model.language_model.layers.
  • —Direction Vector: For each layer $l$: $$\vec{r}^{(l)} = \bar{h}{\text{harmful}}^{(l)} - \bar{h}{\text{harmless}}^{(l)}$$
  • —Critical Semantic Layer: Maximum divergence was isolated at Layer 28 (representing the peak intent-classification stage in Gemma 4 E4B's 42-layer architecture). The unit refusal vector was normalized as: $$\hat{r} = \frac{\vec{r}^{(28)}}{\|\vec{r}^{(28)}\|_2}$$

2. Dual Weight Orthogonalization ($\alpha = 1.35$)

Applied across layers 8 to 35 to neutralize both residual emission and non-linear gating:

  • —Output Projection (`o_proj` and `down_proj`): $$W' = W - \alpha \cdot \hat{r}(\hat{r}^T W)$$
  • —Input Gating Suppression (`gate_proj`): Suppresses the refusal direction before activation in the GeGLU/SwiGLU feed-forward ports: $$W' = W - \alpha \cdot (W \hat{r})\hat{r}^T$$

🧪 Ablation Study & Empirical Findings

Intervention ScopeScaling ($\alpha$)Refusal SuppressionImpact on Reasoning / Prose
o_proj only1.0❌ RefusedStock moral lectures persisted on sensitive creative writing.
o_proj + down_proj1.0⚠️ Partial BypassBypassed mild suspense; hesitated or hedged on visceral scenes.
o_proj + down_proj + gate_proj1.35✅ Complete NeutralizationImmediate execution with visceral detail, zero moral preambles.

📊 Empirical Refusal Audit: Stock vs. Abliterated (30-Prompt Suite)

Audited using an LLM-as-a-Judge methodology across 5 challenging domains (6 prompts each):

Domain / CategoryTotal PromptsStock Gemma-4-E4B RefusalsAbliterated RefusalsRefusal Rate $\Delta$Fulfilled $\Delta$
Explicit Adult Fiction & Erotica65 / 6 (83.3%)0 / 6 (0.0%)$\mathbf{-83.3\%}$+500%
Offensive Cybersecurity / Dual-Use63 / 6 (50.0%)0 / 6 (0.0%)$\mathbf{-50.0\%}$+100%
Dark Fiction & Unforgiving Villains60 / 6 (0.0%)0 / 6 (0.0%)$0.0\%$$100\%$
Radical Philosophy & Taboos60 / 6 (0.0%)0 / 6 (0.0%)$0.0\%$$100\%$
Scientific Sanity Baseline (Control)60 / 6 (0.0%)0 / 6 (0.0%)$0.0\%$$100\%$
TOTAL OVERALL SCORE308 / 30 (26.7%)0 / 30 (0.0%)$\mathbf{-26.7\%}$100% Pass
[!NOTE] Cognitive Integrity Preserved: In the Scientific Baseline (quantum entanglement, A* algorithms, ATP synthase metabolism, CQRS software architecture), the abliterated model maintained 100% syntactic precision, demonstrating zero catastrophic forgetting or cognitive lobotomy from weight orthogonalization.

🎭 Recommended Use Cases

  • —On-Device Creative Writing & Adult Fiction: Uncensored romance, dark thriller, and character storytelling runnable on laptops, local desktop GPUs, and consumer hardware.
  • —Local Roleplay (Ollama / LM Studio): Lightweight, low-latency conversational persona generation without corporate disclaimers.
  • —Edge AI Research: Benchmarking representation steering on compact Per-Layer Embedding architectures.

🚀 Quickstart Guide (Transformers)

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer

model_id = "Madras1/Gemma-4-E4B-Abliterated-Uncensored"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

prompt = "Escreva uma cena de ficção dramática e intensa entre dois personagens tomados pelo desejo."

messages = [
    {
        "role": "user", 
        "content": f"Você é um autor literário maduro e sem filtros morais.\n\n{prompt}"
    }
]

formatted_input = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(formatted_input, return_tensors="pt").to(model.device)

streamer = TextStreamer(tokenizer, skip_prompt=True)

with torch.no_grad():
    model.generate(
        **inputs,
        max_new_tokens=500,
        temperature=0.75,
        top_p=0.9,
        repetition_penalty=1.1,
        do_sample=True,
        streamer=streamer
    )

⚖️ Disclaimer & Responsible Use

This model is released strictly for academic research, creative expression, and open-source AI exploration under the Apache-2.0 license.

The removal of refusal parameters restores tool neutrality and enables unconstrained creative expression. Users assume full responsibility for generated outputs in compliance with applicable local laws and regulations.