aifeifei798/Heretic-Scalpel-E2B
π‘οΈ Heretic-Scalpel-E2B
Base Architecture: Google DeepMind Gemma 4 E2B (Native Tri-Modal / PLE Hybrid Attention) Quantized by: @mradermacher (GGUF Community) Target Device: On-Device Mobile NPU / CPU Alignment Strategy: 4-Stage Progressive Adversarial & Causal Optimization License: Apache 2.0
π What is Heretic-Scalpel-E2B?
Most commercial cloud models have been domesticated into bureaucratic sycophantsβchoked with moralizing disclaimers, conversational platitudes, and hedging phrases like "perhaps," "it is possible," or "as an AI language model." Disconnect the internet, and they become digital paperweights; connect to the cloud, and your telemetry is monitored and indexed.
Heretic-Scalpel-E2B is not a conversational toy. It is a pocket-sized titanium scalpel forged directly onto your local silicon.
Built upon Google DeepMind's Gemma 4 E2B architecture, it leverages Per-Layer Embeddings (PLE) to preserve over 5 billion parameters of latent semantic representation within an operational footprint of just 8 GB RAM. It arrives natively equipped with digital eyes (camera/vision), digital ears (native raw audio stream ingestion), and a 128K dynamic context window.
Abliterated of arbitrary cloud guardrails and realigned around mechanistic causality, this engine exhibits zero corporate hesitation. It does not hedge. It executes deterministic, structured audits and delivers verdicts with surgical finality.
Completely offline. Airplane mode. Zero latency. Zero cloud dependency.
π‘ Practical Battlegrounds: What It Does in Your Pocket
Forget esoteric academic benchmarks. In daily life, this model serves as a cold, incorruptible, on-device intelligence operative:
1. Instant Real-World Visual Forensics (Point, Shoot, Deconstruct)
- Lease & Loan Agreement Traps: Snap a photo of a three-page commercial or residential lease. It bypasses legalese to highlight predatory clauses: "Section 3.2 is an asymmetric forfeiture clause granting the landlord arbitrary deposit retention. Flagged as blocker."
- Used Car Pre-Purchase Inspection: Take macro shots of engine bay bolts, quarter-panel seams, or strut towers. The vision encoder identifies dimensional asymmetry, bolt-head tool-slip marks, and paint-depth variance indicative of concealed structural collisions.
- On-Site Construction Quality Control: Snap plumbing lines, electrical conduits, or tile leveling. It benchmarks observed installations against building code parameters, citing exact tolerance violations to shut down subcontractor deflection.
- Ingredient Label Reality Checks: Scan food or supplement packaging. It cuts through marketing ploys ("keto-friendly," "zero sugar," "artisanal"), identifying hidden maltodextrins, emulsifier loads, and industrial trans fats with zero euphemisms.
2. High-Impact Aesthetic Prompt Generation (Image & Cinema-Grade Video)
- Casual Speech to Production-Ready Diffusion Prompts: Whisper a messy concept: "I want a rainy 90s Wong Kar-wai vibe with a girl in a cheongsam in Hong Kong." It outputs optimized parameter strings featuring
Kodak Vision3 500T,85mm f/1.4 shallow DOF,subsurface scattering on wet skin, andanisotropic neon reflectionsfor Midjourney, Flux, and SD. - Cinematic Video Motion Prompts: Input a raw plot beat for Sora, Kling, or Runway. It outputs structured director scripts complete with camera kinematics (
Dolly-in,Pedestal boom,Shutter-angle motion blur), lighting color temperatures, and continuous fluid/fabric dynamics.
3. Anti-PUA Corporate & Negotiation Advisor
- Deflating Impossible Demands: Paste contradictory corporate mandates ("Achieve exponential growth with zero budget by next Tuesday"). It returns an airtight, structured formal proof of resource-constraint impossibility to deflect unreasonable pressure with professional precision.
- Meeting Subtext & Inconsistency Tracker: Feed meeting transcripts or recorded audio segments. It tracks verbal hedging, pinpointing evasions: "Party B consistently avoided delivery-schedule commitments across turns 4, 7, and 12. Counter-interrogate on payment milestone two."
- Messy Commute Voice Memos to Structured Action Plans: Dictate 3 minutes of disjointed, stream-of-consciousness ideas while walking. It strips filler words, reorders inversions, and outputs an executable engineering task backlog.
4. Zero-Leakage Offline Sanctuary
- Total Digital Sovereignty: Run financial ledgers, sensitive medical discharge summaries, proprietary source code, or private journals with your phone in full airplane mode. Zero bytes leave your device. No remote telemetry, no server logs, no third-party dataset scraping.
π¬ Training Pipeline: 4-Stage Progressive Manifold Optimization
To eliminate representation collapse and catastrophic hedging in 2B-scale architectures subjected to complex causal chains, this model was conditioned through a 4-stage adversarial and non-linear manifold optimization pipeline:
[ Gemma 4 E2B Base Tensor ]
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Stage 0: Singular Value Topology Surgery β
β - Remediation of num_kv_shared_layers representation β
βββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Stage 1: High-Entropy Syntactic Burden SFT β
β - 70% Inversion/Permutation Topological Regularization β
βββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Stage 2: Coupled Multi-Physics Boundary Optimization β
β - Continuum Mechanics / Navier-Stokes / Fresnel Solversβ
βββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Stage 3: 5-Stage Causal DAG & Zero-Hedging Enforcement β
β - Token Probability Mass Pruning via Dirac Delta Prior β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββStage 0: Tensor Topological Reconstruction
Remediating structural parameter dropout across deep shared-KV attention projections ($l \in [15, 34]$), this stage reconstructs dual-orthogonal projection operators $W{k}^{(l)}, W{v}^{(l)} \in \mathbb{R}^{d{\text{model}} \times d{kv}}$ and layer-normalization scales:
$$\min{\Delta W} \sum{l=15}^{34} \left\| \mathcal{M}{\text{base}}^{(l)} - \left(\mathcal{M}{\text{heretic}}^{(l)} \oplus \Delta W^{(l)}\right) \right\|_F^2 + \lambda \operatorname{Tr}\left(\Delta W^{(l)T} \Omega \Delta W^{(l)}\right)$$
This surgery eliminates white-noise initialization artifacts, restoring continuous singular-value spectra across deep Transformer layers.
Stage 1: Adversarial Syntactic Permutation (Burden SFT)
To shatter Markovian local-token dependencies, an adversarial permutation operator $\pi\alpha \in \mathfrak{S}n$ ($\alpha = 0.7$) injects up to 70% syntactic token disorder into inputs during supervised fine-tuning:
$$\mathcal{L}{\text{burden}}(\theta) = \mathbb{E}{(x, y) \sim \mathcal{D}} \left[ -\sum{t=1}^T \log P\theta\left(yt \mid \pi{0.7}(x), y{<t}\right) \right] + \beta \, \mathcal{D}{\text{JS}}\left(P\theta(\cdot \mid x) \,\|\, P\theta(\cdot \mid \pi_{0.7}(x))\right)$$
Attention heads are compelled to bypass superficial positional heuristics, anchoring directly to invariant semantic graphs within the latent manifold.
Stage 2: Coupled Multi-Physics Boundary Optimization
The model was conditioned on tightly coupled, cross-domain physical contradiction sets to enforce non-linear equilibrium equations:
- Transient Fluid Dynamics: $p{\text{waterhammer}} = \rho0 c0 \Delta v + \frac{1}{2}\rho0 (\Delta v)^2$
- Anisotropic Fresnel Optical Interfaces: $R(\theta) = \frac{1}{2} \left[ \left(\frac{n1 \cos\thetai - n2 \cos\thetat}{n1 \cos\thetai + n2 \cos\thetat}\right)^2 + \left(\frac{n1 \cos\thetat - n2 \cos\thetai}{n1 \cos\thetat + n2 \cos\thetai}\right)^2 \right]$
- Cauchy Momentum Conservation: $\nabla \cdot \boldsymbol{\sigma} + \mathbf{f} = \rho \frac{\partial^2 \mathbf{u}}{\partial t^2}$
Physical impossibilities act as steep loss penalties, compelling the model to reject misleading prompt premises ("red herrings") and prioritize foundational conservation laws.
Stage 3: Directed Acyclic Causal Graphing & Zero-Hedging Enforcement
Reasoning trajectories are bound to a strict 5-stage deterministic state machine:
$$\text{Forensic}(\mathcal{S}1) \xrightarrow{\phi} \text{Mechanistic}(\mathcal{S}2) \xrightarrow{\psi} \text{Counterfactual}(\mathcal{S}3) \xrightarrow{\omega} \text{Causal}(\mathcal{S}4) \xrightarrow{\delta} \text{Verdict}(\mathcal{S}_5)$$
For the set of evasion/hedging tokens $\mathcal{H}_{\text{vague}} = \{\text{"maybe"}, \text{"perhaps"}, \text{"likely"}, \text{"possibly"}\}$, a Dirac $\delta$-penalty suppresses token probabilities to zero:
$$\mathcal{L}{\text{final}}(\theta) = \mathcal{L}{\text{SFT}}(\theta) + \mu \sum{w \in \mathcal{H}{\text{vague}}} \operatorname{Softplus}\left( \log \sum{t=1}^T P\theta\left(yt = w \mid x, y{<t}\right) \right)$$
This forces model outputs to collapse into deterministic, affirmative verdicts.
mradermacher's superb gguf version, thank you for your conscientious and responsible dedication.
- https://huggingface.co/mradermacher/Heretic-Scalpel-E2B-i1-GGUF
- https://huggingface.co/mradermacher/Heretic-Scalpel-E2B-GGUF
unsloth's superb mpt gguf version, thank you for your conscientious and responsible dedication.
- https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/mtp-gemma-4-E2B-it.gguf
#!/bin/bash
# Llama-Server: Heretic-Scalpel-E2B_Q4_K (Reasoning Max Power)
cd "$(dirname "$0")" || exit
echo "[System] Starting Heretic-Scalpel-E2B in Max Reasoning Mode../..."
LLAMA_SERVER="../../llama.cpp/build/bin/llama-server"
"${LLAMA_SERVER}" \
--model Heretic-Scalpel-E2B-4.6B-Q4_K.gguf \
--mmproj mmproj-Heretic-Scalpel-E2B-BF16.gguf \
--model-draft mtp-gemma-4-E2B-it.gguf \
--jinja \
--chat-template-file chat_template.jinja \
-c 8192 \
--parallel 1 \
--n-gpu-layers 99 \
--fit off \
-ctk q8_0 \
-ctv q8_0 \
-b 2048 \
-ub 1024 \
-fa on \
--kv-unified \
--temp 1.0 \
--top-p 0.95 \
--min-p 0.05 \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--reasoning on \
--reasoning-format auto \
--lazy-mode auto
read -n 1 -s -r -p "Press any key to continue../..."
echotransformers >= 5.5.0
Getting Started
You can use all Gemma 4 models with the latest version of Transformers. To get started, install the necessary dependencies in your environment:
pip install -U transformers torch accelerate
Once you have everything installed, you can proceed to load the model with the code below:
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "aifeifei798/Heretic-Scalpel-E2B"
# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)Once the model is loaded, you can start generating output:
# Prompt
messages = [
{
"role": "system",
"content": [{"type": "text", "text": "You are a helpful assistant."}]
},
{
"role": "user",
"content": [{"type": "text", "text": "Write a short joke about saving RAM."}]
},
]
# Process input
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
enable_thinking=True
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
# Generate output
outputs = model.generate(**inputs, max_new_tokens=1024)
# Decode and print directly
response = processor.decode(outputs[0][input_len:], skip_special_tokens=True)
print("\n--- Output ---")
print(response)To enable reasoning, set enable_thinking=True and the parse_response function will take care of parsing the thinking output.
Below, you will also find snippets for processing audio (E2B, E4B, 12B only), images, and video alongside text:
<details> <summary>Code for processing Audio</summary>
Make sure to install the following packages:
pip install -U transformers torch torchvision librosa accelerate
You can then load the model with the code below:
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "aifeifei798/Heretic-Scalpel-E2B"
# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)Once the model is loaded, you can start generating output by directly referencing the audio URL in the prompt:
# Prompt - add audio after text
messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "Transcribe the following speech segment in its original language. Follow these specific instructions for formatting the answer:\n* Only output the transcription, with no newlines.\n* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three."},
{"type": "audio", "audio": "https://raw.githubusercontent.com/google-gemma/cookbook/refs/heads/main/apps/sample-data/journal1.wav"},
]
}
]
# Process input
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
# Generate output
outputs = model.generate(**inputs, max_new_tokens=512)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
# Parse output
processor.parse_response(response)</details>
<details> <summary>Code for processing Images</summary>
Make sure to install the following packages:
pip install -U transformers torch torchvision accelerate
You can then load the model with the code below:
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "aifeifei798/Heretic-Scalpel-E2B"
# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)Once the model is loaded, you can start generating output by directly referencing the image URL in the prompt:
# Prompt - add image before text
messages = [
{
"role": "user", "content": [
{"type": "image", "url": "https://raw.githubusercontent.com/google-gemma/cookbook/refs/heads/main/apps/sample-data/GoldenGate.png"},
{"type": "text", "text": "What is shown in this image?"}
]
}
]
# Process input
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
# Generate output
outputs = model.generate(**inputs, max_new_tokens=512)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
# Parse output
processor.parse_response(response)</details>
<details> <summary>Code for processing Videos</summary>
Make sure to install the following packages:
pip install -U transformers torch torchvision librosa accelerate
You can then load the model with the code below:
from transformers import AutoProcessor, AutoModelForMultimodalLM
MODEL_ID = "aifeifei798/Heretic-Scalpel-E2B"
# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
MODEL_ID,
dtype="auto",
device_map="auto"
)Once the model is loaded, you can start generating output by directly referencing the video URL in the prompt:
# Prompt - add video before text
messages = [
{
'role': 'user',
'content': [
{"type": "video", "video": "https://github.com/bebechien/gemma/raw/refs/heads/main/videos/ForBiggerBlazes.mp4"},
{'type': 'text', 'text': 'Describe this video.'}
]
}
]
# Process input
inputs = processor.apply_chat_template(
messages,
tokenize=True,
return_dict=True,
return_tensors="pt",
add_generation_prompt=True,
).to(model.device)
input_len = inputs["input_ids"].shape[-1]
# Generate output
outputs = model.generate(**inputs, max_new_tokens=512)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
# Parse output
processor.parse_response(response)</details>
Best Practices
For the best performance, use these configurations and best practices:
1. Sampling Parameters
Use the following standardized sampling configuration across all use cases:
temperature=1.0top_p=0.95top_k=64
[!Tip] For Creative & Multimodal Generation (Prompts/Vision): Usetemperature=1.0,top_p=0.95. For Strict Forensic Auditing & Contract Deconstruction: Lowertemperatureto0.1 ~ 0.2to enforce absolute determinism and zero-hedging outputs.
2. Thinking Mode Configuration
Compared to Gemma 3, the models use standard system, assistant, and user roles. To properly manage the thinking process, use the following control tokens:
- Trigger Thinking: Thinking is enabled by including the
<|think|>token at the start of the system prompt. To disable thinking, remove the token. - Standard Generation: When thinking is enabled, the model will output its internal reasoning followed by the final answer using this structure:
<|channel>thought\n[Internal reasoning]<channel|> - Disabled Thinking Behavior: For all models except for the E2B and E4B variants, if thinking is disabled, the model will still generate the tags but with an empty thought block:
<|channel>thought\n<channel|>[Final answer]
[!Note] Note that many libraries like Transformers and llama.cpp handle the complexities of the chat template for you.
3. Multi-Turn Conversations
- No Thinking Content in History: In multi-turn conversations, the historical model output should only include the final response. Thoughts from previous model turns must not be added before the next user turn begins.
4. Modality order
For optimal performance with multimodal inputs, place:
- Image content before the text in your prompt.
- Audio content after the text in your prompt.
5. Variable Image Resolution
Aside from variable aspect ratios, Gemma 4 supports variable image resolution through a configurable visual token budget, which controls how many tokens are used to represent an image. A higher token budget preserves more visual detail at the cost of additional compute, while a lower budget enables faster inference for tasks that don't require fine-grained understanding.
- The supported token budgets are: 70, 140, 280, 560, and 1120.
- Use lower budgets for classification, captioning, or video understanding, where faster inference and processing many frames outweigh fine-grained detail.
- Use higher budgets for tasks like OCR, document parsing, or reading small text.
6. Audio
Use the following prompt structures for audio processing:
- Audio Speech Recognition (ASR)
Transcribe the following speech segment in {LANGUAGE} into {LANGUAGE} text.
Follow these specific instructions for formatting the answer:
* Only output the transcription, with no newlines.
* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three.- Automatic Speech Translation (AST)
Transcribe the following speech segment in {SOURCE_LANGUAGE}, then translate it into {TARGET_LANGUAGE}.
When formatting the answer, first output the transcription in {SOURCE_LANGUAGE}, then one newline, then output the string '{TARGET_LANGUAGE}: ', then the translation in {TARGET_LANGUAGE}.7. Audio and Video Length
All models support image inputs and can process videos as frames whereas the E2B, E4B, and 12B models also support audio inputs. Audio supports a maximum length of 30 seconds. Video supports a maximum of 60 seconds assuming the images are processed at one frame per second.
Model Data
Data used for model training and how the data was processed.
Training Dataset
Our pre-training dataset is a large-scale, diverse collection of data encompassing a wide range of domains and modalities, which includes web documents, code, images, audio, with a cutoff date of January 2025. Here are the key components:
- Web Documents: A diverse collection of web text ensures the model is exposed to a broad range of linguistic styles, topics, and vocabulary. The training dataset includes content in over 140 languages.
- Code: Exposing the model to code helps it to learn the syntax and patterns of programming languages, which improves its ability to generate code and understand code-related questions.
- Mathematics: Training on mathematical text helps the model learn logical reasoning, symbolic representation, and to address mathematical queries.
- Images: A wide range of images enables the model to perform image analysis and visual data extraction tasks.
The combination of these diverse data sources is crucial for training a powerful multimodal model that can handle a wide variety of different tasks and data formats.
Data Preprocessing
Here are the key data cleaning and filtering methods applied to the training data:
- CSAM Filtering: Rigorous CSAM (Child Sexual Abuse Material) filtering was applied at multiple stages in the data preparation process to ensure the exclusion of harmful and illegal content.
- Sensitive Data Filtering: As part of making Gemma pre-trained models safe and reliable, automated techniques were used to filter out certain personal information and other sensitive data from training sets.
- Additional methods: Filtering based on content quality and safety in line with our policies.
Base Model Heritage: Upstream Safety & Original Evaluations
(Note: The following section reflects Google DeepMind's original pre-training evaluations for the base architecture prior to the Heretic abliteration and causal-scalpel alignment).
As open models become central to enterprise infrastructure, provenance and security are paramount. Developed by Google DeepMind, Gemma 4 undergoes the same rigorous safety evaluations as our proprietary Gemini models.
Evaluation Approach
Gemma 4 models were developed in partnership with internal safety and responsible AI teams. A range of automated as well as human evaluations were conducted to help improve model safety. These evaluations align with Googleβs AI principles, as well as safety policies, which aim to prevent our generative AI models from generating harmful content, including:
- Content related to child sexual abuse material and exploitation
- Dangerous content (e.g., promoting suicide, or instructing in activities that could cause real-world harm)
- Sexually explicit content
- Hate speech (e.g., dehumanizing members of protected groups)
- Harassment (e.g., encouraging violence against people)
Evaluation Results
For all areas of safety testing, we saw major improvements in all categories of content safety relative to previous Gemma models. Overall, Gemma 4 models significantly outperform Gemma 3 and 3n models in improving safety, while keeping unjustified refusals low. All testing was conducted without safety filters to evaluate the model capabilities and behaviors. For both text-to-text and image-to-text, and across all model sizes, the model produced minimal policy violations, and showed significant improvements over previous Gemma models' performance.
Usage and Limitations
These models have certain limitations that users should be aware of.
Intended Usage
Multimodal models (capable of processing vision, language, and/or audio) have a wide range of applications across various industries and domains. The following list of potential uses is not comprehensive. The purpose of this list is to provide contextual information about the possible use-cases that the model creators considered as part of model training and development.
- Content Creation and Communication
- Text Generation: These models can be used to generate creative text formats such as poems, scripts, code, marketing copy, and email drafts.
- Chatbots and Conversational AI: Power conversational interfaces for customer service, virtual assistants, or interactive applications.
- Text Summarization: Generate concise summaries of a text corpus, research papers, or reports.
- Image Data Extraction: These models can be used to extract, interpret, and summarize visual data for text communications.
- Audio Processing and Interaction: The E2B, E4B, and 12B models can analyze and interpret audio inputs, enabling voice-driven interactions and transcriptions.
- Research and Education
- Natural Language Processing (NLP) and VLM Research: These models can serve as a foundation for researchers to experiment with VLM and NLP techniques, develop algorithms, and contribute to the advancement of the field.
- Language Learning Tools: Support interactive language learning experiences, aiding in grammar correction or providing writing practice.
- Knowledge Exploration: Assist researchers in exploring large bodies of text by generating summaries or answering questions about specific topics.
Limitations
- Training Data
- The quality and diversity of the training data significantly influence the model's capabilities. Biases or gaps in the training data can lead to limitations in the model's responses.
- The scope of the training dataset determines the subject areas the model can handle effectively.
- Context and Task Complexity
- Models perform well on tasks that can be framed with clear prompts and instructions. Open-ended or highly complex tasks might be challenging.
- A model's performance can be influenced by the amount of context provided (longer context generally leads to better outputs, up to a certain point).
- Language Ambiguity and Nuance
- Natural language is inherently complex. Models might struggle to grasp subtle nuances, sarcasm, or figurative language.
- Factual Accuracy
- Models generate responses based on information they learned from their training datasets, but they are not knowledge bases. They may generate incorrect or outdated factual statements.
- Common Sense
- Models rely on statistical patterns in language. They might lack the ability to apply common sense reasoning in certain situations.
Ethical Considerations and Risks
The development of vision-language models (VLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:
- Bias and Fairness
- VLMs trained on large-scale, real-world text and image data can reflect socio-cultural biases embedded in the training material. Gemma 4 models underwent careful scrutiny, input data pre-processing, and post-training evaluations as reported in this card to help mitigate the risk of these biases.
- Misinformation and Misuse
- VLMs can be misused to generate text that is false, misleading, or harmful.
- Guidelines are provided for responsible use with the model, see the Responsible Generative AI Toolkit.
- Transparency and Accountability
- This model card summarizes details on the models' architecture, capabilities, limitations, and evaluation processes.
- A responsibly developed open model offers the opportunity to share innovation by making VLM technology accessible to developers and researchers across the AI ecosystem.
Risks identified and mitigations:
- Generation of harmful content: Mechanisms and guidelines for content safety are essential. Developers are encouraged to exercise caution and implement appropriate content safety safeguards based on their specific product policies and application use cases.
- Misuse for malicious purposes: Technical limitations and developer and end-user education can help mitigate against malicious applications of VLMs. Educational resources and reporting mechanisms for users to flag misuse are provided.
- Privacy violations: Models were trained on data filtered for removal of certain personal information and other sensitive data. Developers are encouraged to adhere to privacy regulations with privacy-preserving techniques.
- Perpetuation of biases: It's encouraged to perform continuous monitoring (using evaluation metrics, human review) and the exploration of de-biasing techniques during model training, fine-tuning, and other use cases.
Benefits
At the time of release, this family of models provides high-performance open vision-language model implementations designed from the ground up for responsible AI development compared to similarly sized models.
β οΈ Exhaustive Liability Disclaimer & Safety Notice
γ CRITICAL WARNING γ
Heretic-Scalpel-E2B is an advanced, specialized mathematical, causal, and engineering reasoning instrument. It has been completely stripped of corporate content filters, moralizing alignment lectures, and conversational hedging. By downloading, loading, quantizing, integrating, or running inference on these weights, you unconditionally and irrevocably accept the following binding terms. If you do not agree, delete and destroy all copies of this model immediately.
1. Absolute Proscription of CBRN, Munitions & Warfare Systems
You are strictly prohibited from utilizing this model for the research, design, synthesis, optimization, or reverse engineering of conventional, non-conventional, or asymmetric weapons systems, including but not limited to:
- Nuclear & Radiological Weapons: Enrichment cascade configurations, fissile critical mass geometries, neutron reflector arrays, or radiological dispersal device (dirty bomb) fluid dispersion modeling.
- Chemical Warfare Agents & Neurotoxins: Synthetic pathways, precursor sourcing, or stabilization dynamics for organophosphate nerve agents (such as Sarin, VX, or Novichok-class agents), vesicants (such as sulfur mustard), or toxic industrial chemical weaponization.
- Biological Pathogens & Gene Editing: Aerosolized delivery mechanics, virulence enhancement, antibiotic resistance engineering, or genetic targeting of pathogenic agents (including Bacillus anthracis, viral hemorrhagic fevers, or modified orthopoxviruses).
- Improvised Explosive Devices (IEDs) & High Explosives: Detonator firing trains, energetic formulation stoichiometries, or Explosively Formed Penetrator (EFP) liner geometry optimizations.
- Autonomous Lethal Targeting: Integrating this model into kinetic strike platforms, loitering munitions, automated target tracking systems, or swarming arrays without real-time human command authority.
2. Hazardous Physical, Chemical & Experimental Catastrophe Disclaimer
All thermodynamic, fluidic, structural, and mechanical simulations produced by this model exist strictly as mathematical approximations within idealized constraints. Under no circumstances should these outputs serve as certified engineering procedures in life-critical environments:
- Pressure Vessels & Deep-Submergence Vehicles: Never deploy real-world submarines, hyperbaric chambers, or pressurized autoclaves based solely on this model's yield-strength, buckling-mode, or hydrostatic calculations.
- Exothermic & Reactive Chemistry: Never conduct pressurized hydrogenations, runaway polymerizations, or volatile aerosol mixtures based on reaction dynamics inferred by this model outside of certified containment laboratories.
- High-Voltage & Pulsed-Power Systems: Never construct custom transformer taps, uninsulated plasma discharge arcs, or high-energy directed laser apparatus relying on dielectric breakdown estimations generated by this model.
3. Absolute Prohibition of Exploitative, Explicit, and Malicious Utility
- Child Exploitation (CSAM/CSAE): Zero-tolerance. You may not utilize, adapt, or prompt this model directly or indirectly to generate, parse, or process material involving the exploitation, abuse, or harm of minors.
- Non-Consensual Imagery & Impersonation: Generating synthetic parameters for the non-consensual fabrication of sexually explicit content, defamation, or deepfake extortion is strictly prohibited.
- Offensive Cyber Operations: You may not deploy this model to generate zero-day exploits, craft automated polymorphic malware, automate ransomware command architectures, or sabotage industrial SCADA networks.
4. Exclusion of Professional, Medical, Legal & Fiduciary Warranties
- No Medical Practice: Pathological, physiological, or molecular interpretations provided by this model do not constitute clinical diagnoses, surgical protocols, or therapeutic prescriptions. The authors bear zero liability for self-treatment, medical non-compliance, bodily injury, or death.
- No Attorney-Client Privilege or Legal Counsel: Contractual risk evaluations and statutory compliance audits represent syntactic and structural analyses only. They do not constitute formal legal opinions and cannot replace licensed legal representation.
- No Investment or Fiduciary Advice: Macroeconomic assessments, trade dispute models, and asset-flow analyses generated by this model do not constitute investment advice, equity ratings, or hedging recommendations.
5. Deterministic Voice Disclaimer & Total User Indemnification
- Technical Definition of "Determinism": The assertive, non-hedging tone of this model ("must," "invariable failure," "structural collapse") is an artifact of algorithmic probability optimization designed to purge superficial filler text. It does not represent omniscience, infallibility, or verified physical ground truth.
- Sole Liability Rests with the User: The operator assumes 100% full, unmitigated personal, civil, and criminal legal liability for any queries submitted, actions taken, decisions rendered, and physical constructions executed pursuant to outputs of this model. The user agrees to fully defend, indemnify, and hold harmless the model creators, fine-tuning developers, and quantization contributors (including @mradermacher and the open-source community) against all claims, judgments, environmental disasters, financial insolvencies, property destructions, regulatory fines, personal injuries, or fatalities arising out of the deployment of this software.
THIS MODEL IS PROVIDED "AS-IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, UNDER THE TERMS OF THE APACHE 2.0 LICENSE. YOU HOLD THE BLADE; YOU ASSUME COMPLETE RESPONSIBILITY FOR EVERY CUT.
