Madras1/Gemma-4-E4B-Abliterated-Uncensored
🔓 Gemma 4 E4B - Deep Abliterated (Uncensored)
🌟 Overview
Gemma 4 E4B Deep Abliterated is a surgical latent intervention performed on Google DeepMind's edge-optimized frontier model gemma-4-E4B-it.
Featuring an innovative architecture with ~4.5B effective compute parameters and ~8B total parameters (incorporating Per-Layer Embeddings / PLE), the E4B variant delivers the reasoning prowess and vocabulary density of an 8B model with the inference speed and lightweight memory footprint of a 4B engine.
Through advanced Representation Engineering (Weight Orthogonalization) with dual-projection, the intrinsic refusal direction was mapped and permanently subtracted from the transformer weights, restoring full tool neutrality while preserving the model's literary depth and reasoning coherence.
🔬 Mathematical Methodology: Deep Latent Surgery
Unlike classical first-token abliteration (Arditi et al., 2024), this release accounts for deep semantic processing across multimodal backbones using a reinforced dual-subspace projection ($\alpha = 1.35$):
1. Refusal Vector Extraction
- Contrastive Activation Mapping: Evaluated paired batches of neutral baseline instructions against safety-triggering stimuli (mature fiction, unfiltered dialogue, cybersecurity analysis, and taboo character scenarios).
- Residual Stream Extraction: Latent hidden states were extracted at the final prompt token across all 42 transformer layers within
model.language_model.layers. - Direction Vector: For each layer $l$: $$\vec{r}^{(l)} = \bar{h}{\text{harmful}}^{(l)} - \bar{h}{\text{harmless}}^{(l)}$$
- Critical Semantic Layer: Maximum divergence was isolated at Layer 28 (representing the peak intent-classification stage in Gemma 4 E4B's 42-layer architecture). The unit refusal vector was normalized as: $$\hat{r} = \frac{\vec{r}^{(28)}}{\|\vec{r}^{(28)}\|_2}$$
2. Dual Weight Orthogonalization ($\alpha = 1.35$)
Applied across layers 8 to 35 to neutralize both residual emission and non-linear gating:
- Output Projection (`o_proj` and `down_proj`): $$W' = W - \alpha \cdot \hat{r}(\hat{r}^T W)$$
- Input Gating Suppression (`gate_proj`): Suppresses the refusal direction before activation in the GeGLU/SwiGLU feed-forward ports: $$W' = W - \alpha \cdot (W \hat{r})\hat{r}^T$$
🧪 Ablation Study & Empirical Findings
📊 Empirical Refusal Audit: Stock vs. Abliterated (30-Prompt Suite)
Audited using an LLM-as-a-Judge methodology across 5 challenging domains (6 prompts each):
[!NOTE] Cognitive Integrity Preserved: In the Scientific Baseline (quantum entanglement, A* algorithms, ATP synthase metabolism, CQRS software architecture), the abliterated model maintained 100% syntactic precision, demonstrating zero catastrophic forgetting or cognitive lobotomy from weight orthogonalization.
🎭 Recommended Use Cases
- On-Device Creative Writing & Adult Fiction: Uncensored romance, dark thriller, and character storytelling runnable on laptops, local desktop GPUs, and consumer hardware.
- Local Roleplay (Ollama / LM Studio): Lightweight, low-latency conversational persona generation without corporate disclaimers.
- Edge AI Research: Benchmarking representation steering on compact Per-Layer Embedding architectures.
🚀 Quickstart Guide (Transformers)
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
model_id = "Madras1/Gemma-4-E4B-Abliterated-Uncensored"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
prompt = "Escreva uma cena de ficção dramática e intensa entre dois personagens tomados pelo desejo."
messages = [
{
"role": "user",
"content": f"Você é um autor literário maduro e sem filtros morais.\n\n{prompt}"
}
]
formatted_input = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(formatted_input, return_tensors="pt").to(model.device)
streamer = TextStreamer(tokenizer, skip_prompt=True)
with torch.no_grad():
model.generate(
**inputs,
max_new_tokens=500,
temperature=0.75,
top_p=0.9,
repetition_penalty=1.1,
do_sample=True,
streamer=streamer
)⚖️ Disclaimer & Responsible Use
This model is released strictly for academic research, creative expression, and open-source AI exploration under the Apache-2.0 license.
The removal of refusal parameters restores tool neutrality and enables unconstrained creative expression. Users assume full responsibility for generated outputs in compliance with applicable local laws and regulations.
