coder3101/gemma-4-26B-A4B-it-qat-q4_0-unquantized-heretic
1151
1---2library_name: transformers3license: apache-2.04license_link: https://ai.google.dev/gemma/docs/gemma_4_license5pipeline_tag: image-text-to-text6base_model:7- google/gemma-4-26B-A4B-it-qat-q4_0-unquantized8tags:9- heretic10- uncensored11- decensored12- abliterated13- ara14---15# This is a decensored version of [google/gemma-4-26B-A4B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized), made using [Heretic](https://github.com/p-e-w/heretic) v1.2.0 with the [Arbitrary-Rank Ablation (ARA)](https://github.com/p-e-w/heretic/pull/211) method (with row-norm preservation)16 17## Abliteration parameters18 19| Parameter | Value |20| :-------- | :---: |21| **start_layer_index** | 12 |22| **end_layer_index** | 21 |23| **preserve_good_behavior_weight** | 0.3106 |24| **steer_bad_behavior_weight** | 0.0066 |25| **overcorrect_relative_weight** | 0.7982 |26| **neighbor_count** | 14 |27 28## Performance29 30| Metric | This model | Original model ([google/gemma-4-26B-A4B-it-qat-q4_0-unquantized](https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized)) |31| :----- | :--------: | :---------------------------: |32| **KL divergence** | 0.0660 | 0 *(by definition)* |33| **Refusals** | 13/100 | 100/100 |34 35-----36 37 38<div align="center">39 <img src=https://ai.google.dev/gemma/images/gemma4_banner.png>40</div>41 42 43<p align="center">44 <a href="https://huggingface.co/collections/google/gemma-4" target="_blank">Hugging Face</a> |45 <a href="https://github.com/google-gemma" target="_blank">GitHub</a> |46 <a href="https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/" target="_blank">Launch Blog</a> |47 <a href="https://ai.google.dev/gemma/docs/core" target="_blank">Documentation</a>48 <br>49 <b>License</b>: <a href="https://ai.google.dev/gemma/docs/gemma_4_license" target="_blank">Apache 2.0</a> | <b>Authors</b>: <a href="https://deepmind.google/models/gemma/" target="_blank">Google DeepMind</a>50</p>51 52> [!Note]53> This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model.54> Four versions of the QAT checkpoints are available:55> * **Unquantized QAT checkpoints** (Q4_0): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models.56> * **GGUF** (Q4_0): Ready-to-deploy formats for broad ecosystem compatibility. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B.57> * **Mobile-optimized** (wNa8o8): A custom schema engineered explicitly for mobile hardware efficiency. It features targeted 2-bit decoding layers, optimized KV caches, and static activations to maximize VRAM savings. Available for Gemma 4 E2B and E4B.58> * **Compressed Tensors** (w4a16): QAT checkpoints serialized in the compressed-tensors format for native, optimized inference with vLLM. Available for Gemma 4 E2B, E4B, 12B, and 31B. 59 60Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages. 61 62Featuring both Dense and Mixture-of-Experts (MoE) architectures, Gemma 4 is well-suited for tasks like text generation, coding, and reasoning. The models are available in five distinct sizes: **E2B**, **E4B**, **12B**, **26B A4B**, and **31B**. Their diverse sizes make them deployable in environments ranging from high-end phones to laptops and servers, democratizing access to state-of-the-art AI.63 64Gemma 4 introduces key **capability and architectural advancements**:65 66* **Reasoning** – All models in the family are designed as highly capable reasoners, with configurable thinking modes.67 68* **Extended Multimodalities** – Processes Text, Image with variable aspect ratio and resolution support (all models), Video, and Audio (featured natively on the E2B, E4B, and 12B models).69 70* **Diverse & Efficient Architectures** – Offers Dense and Mixture-of-Experts (MoE) variants of different sizes for scalable deployment.71 72* **Optimized for On-Device** – Smaller models are specifically designed for efficient local execution on laptops and mobile devices.73 74* **Increased Context Window** – The small models feature a 128K context window, while the medium models support 256K.75 76* **Enhanced Coding & Agentic Capabilities** – Achieves notable improvements in coding benchmarks alongside native function-calling support, powering highly capable autonomous agents.77 78* **Native System Prompt Support** – Gemma 4 introduces native support for the `system` role, enabling more structured and controllable conversations.79 80## **Models Overview**81 82Gemma 4 models are designed to deliver frontier-level performance at each size, targeting deployment scenarios from mobile and edge devices (E2B, E4B) to consumer GPUs and workstations (12B, 26B A4B, 31B). They are well-suited for reasoning, agentic workflows, coding, and multimodal understanding.83 84The models employ a hybrid attention mechanism that interleaves local sliding window attention with full global attention, ensuring the final layer is always global. This hybrid design delivers the processing speed and low memory footprint of a lightweight model without sacrificing the deep awareness required for complex, long-context tasks. To optimize memory for long contexts, global layers feature unified Keys and Values, and apply Proportional RoPE (p-RoPE). 85 86### Dense Models87 88| Property | E2B | E4B | 12B Unified | 31B Dense |89| :---- | :---- | :---- | :---- | :---- |90| **Total Parameters** | 2.3B effective <br> (5.1B with embeddings) | 4.5B effective <br> (8B with embeddings) | 11.95B | 30.7B |91| **Layers** | 35 | 42 | 48 | 60 |92| **Sliding Window** | 512 tokens | 512 tokens | 1024 tokens | 1024 tokens |93| **Context Length** | 128K tokens | 128K tokens | 256K tokens | 256K tokens |94| **Vocabulary Size** | 262K | 262K | 262K | 262K |95| **Supported Modalities** | Text, Image, Audio | Text, Image, Audio | Text, Image, Audio | Text, Image |96| **Vision Encoder Parameters** | *~150M* | *~150M* | - | *~550M* |97| **Audio Encoder Parameters** | *~300M* | *~300M* | - | No Audio |98 99The "E" in E2B and E4B stands for "effective" parameters. The smaller models incorporate Per-Layer Embeddings (PLE) to maximize parameter efficiency in on-device deployments. Rather than adding more layers or parameters to the model, PLE gives each decoder layer its own small embedding for every token. These embedding tables are large but are only used for quick lookups, which is why the effective parameter count is much smaller than the total.100 101The "Unified" in Gemma 4 12B Unified refers to its encoder-free architecture. Other Gemma 4 models use dedicated encoders to process multimodal data before passing it to the LLM. Gemma 4 12B eliminates these encoders entirely, projecting raw image patches and audio waveforms directly into the LLM's embedding space through lightweight linear layers. This unified approach means all modalities flow straight into a single decoder-only transformer, reducing multimodal latency and allowing the entire model to be fine-tuned in one pass.102 103### Mixture-of-Experts (MoE) Model104 105| Property | 26B A4B MoE |106| :---- | :---- |107| **Total Parameters** | 25.2B |108| **Active Parameters** | 3.8B |109| **Layers** | 30 |110| **Sliding Window** | 1024 tokens |111| **Context Length** | 256K tokens |112| **Vocabulary Size** | 262K |113| **Expert Count** | 8 active / 128 total and 1 shared |114| **Supported Modalities** | Text, Image |115| **Vision Encoder Parameters** | *~550M* |116 117The "A" in 26B A4B stands for "active parameters" in contrast to the total number of parameters the model contains. By only activating a 4B subset of parameters during inference, the Mixture-of-Experts model runs much faster than its 26B total might suggest. This makes it an excellent choice for fast inference compared to the dense 31B model since it runs almost as fast as a 4B-parameter model.118 119## **Benchmark Results** 120 121These models were evaluated against a large collection of different datasets and metrics to cover different aspects of text generation. Evaluation results marked in the table are for instruction-tuned models.122 123| | Gemma 4 31B | Gemma 4 26B A4B | Gemma 4 12B Unified | Gemma 4 E4B | Gemma 4 E2B | Gemma 3 27B (no think) |124| :---- | :---- | :---- | :---- | :---- | :---- | :---- |125| MMLU Pro | 85.2% | 82.6% | 77.2% | 69.4% | 60.0% | 67.6% |126| AIME 2026 no tools | 89.2% | 88.3% | 77.5% | 42.5% | 37.5% | 20.8% |127| LiveCodeBench v6 | 80.0% | 77.1% | 72.0% | 52.0% | 44.0% | 29.1% |128| Codeforces ELO | 2150 | 1718 | 1659 | 940 | 633 | 110 |129| GPQA Diamond | 84.3% | 82.3% | 78.8% | 58.6% | 43.4% | 42.4% |130| Tau2 (average over 3) | 76.9% | 68.2% | 69.0% | 42.2% | 24.5% | 16.2% |131| HLE no tools | 19.5% | 8.7% | 5.2% | - | - | - |132| HLE with search | 26.5% | 17.2% | - | - | - | - |133| BigBench Extra Hard | 74.4% | 64.8% | 53.0% | 33.1% | 21.9% | 19.3% |134| MMMLU | 88.4% | 86.3% | 83.4% | 76.6% | 67.4% | 70.7% |135| **Vision** | | | | | | |136| MMMU Pro | 76.9% | 73.8% | 69.1% | 52.6% | 44.2% | 49.7% |137| OmniDocBench 1.5 (average edit distance, lower is better) | 0.131 | 0.149 | 0.164 | 0.181 | 0.290 | 0.365 |138| MATH-Vision | 85.6% | 82.4% | 79.7% | 59.5% | 52.4% | 46.0% |139| MedXPertQA MM | 61.3% | 58.1% | 48.7% | 28.7% | 23.5% | - |140| **Audio** | | | | | | |141| CoVoST | - | - | 38.5<sup>*</sup> | 35.54 | 33.47 | - |142| FLEURS (lower is better) | - | - | 0.069<sup>*</sup> | 0.08 | 0.09 | - |143| **Long Context** | | | | | | |144| MRCR v2 8 needle 128k (average) | 66.4% | 44.1% | 43.4% | 25.4% | 19.1% | 13.5% |145 146<sup>*</sup>Excluding Chinese language.147 148## **Core Capabilities**149 150Gemma 4 models handle a broad range of tasks across text, vision, and audio. Key capabilities include:151 152* **Thinking** – Built-in reasoning mode that lets the model think step-by-step before answering.153* **Long Context** – Context windows of up to 128K tokens (E2B/E4B) and 256K tokens (12B, 26B A4B/31B).154* **Image Understanding** – Object detection, Document/PDF parsing, screen and UI understanding, chart comprehension, OCR (including multilingual), handwriting recognition, and pointing. Images can be processed at variable aspect ratios and resolutions.155* **Video Understanding** – Analyze video by processing sequences of frames.156* **Interleaved Multimodal Input** – Freely mix text and images in any order within a single prompt.157* **Function Calling** – Native support for structured tool use, enabling agentic workflows.158* **Coding** – Code generation, completion, and correction.159* **Multilingual** – Out-of-the-box support for 35+ languages, pre-trained on 140+ languages.160* **Audio** (E2B, E4B, and 12B only) – Automatic speech recognition (ASR) and speech-to-translated-text translation across multiple languages.161 162 163## Getting Started164 165You can use all Gemma 4 models with the latest version of Transformers. To get started, install the necessary dependencies in your environment:166 167`pip install -U transformers torch accelerate`168 169Once you have everything installed, you can proceed to load the model with the code below:170 171```python172from transformers import AutoProcessor, AutoModelForMultimodalLM173 174MODEL_ID = "google/gemma-4-12B-it"175 176# Load model177processor = AutoProcessor.from_pretrained(MODEL_ID)178model = AutoModelForMultimodalLM.from_pretrained(179 MODEL_ID,180 dtype="auto",181 device_map="auto"182)183```184 185Once the model is loaded, you can start generating output:186 187```python188# Prompt189messages = [190 {"role": "system", "content": "You are a helpful assistant."},191 {"role": "user", "content": "Write a short joke about saving RAM."},192]193 194# Process input195inputs = processor.apply_chat_template(196 messages,197 tokenize=True,198 return_dict=True,199 return_tensors="pt",200 add_generation_prompt=True,201 enable_thinking=False202).to(model.device)203input_len = inputs["input_ids"].shape[-1]204 205# Generate output206outputs = model.generate(**inputs, max_new_tokens=1024)207response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)208 209# Parse output210processor.parse_response(response)211```212 213To enable reasoning, set `enable_thinking=True` and the `parse_response` function will take care of parsing the thinking output.214 215Below, you will also find snippets for processing audio (E2B, E4B, 12B only), images, and video alongside text:216 217<details>218<summary>Code for processing Audio</summary>219 220Make sure to install the following packages:221 222`pip install -U transformers torch torchvision librosa accelerate`223 224You can then load the model with the code below:225 226```python227from transformers import AutoProcessor, AutoModelForMultimodalLM228 229MODEL_ID = "google/gemma-4-12B-it"230 231# Load model232processor = AutoProcessor.from_pretrained(MODEL_ID)233model = AutoModelForMultimodalLM.from_pretrained(234 MODEL_ID, 235 dtype="auto", 236 device_map="auto"237)238```239 240Once the model is loaded, you can start generating output by directly referencing the audio URL in the prompt:241 242 243```python244# Prompt - add audio after text245messages = [246 {247 "role": "user",248 "content": [249 {"type": "text", "text": "Transcribe the following speech segment in its original language. Follow these specific instructions for formatting the answer:\n* Only output the transcription, with no newlines.\n* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three."},250 {"type": "audio", "audio": "https://raw.githubusercontent.com/google-gemma/cookbook/refs/heads/main/apps/sample-data/journal1.wav"},251 ]252 }253]254 255# Process input256inputs = processor.apply_chat_template(257 messages,258 tokenize=True,259 return_dict=True,260 return_tensors="pt",261 add_generation_prompt=True,262).to(model.device)263input_len = inputs["input_ids"].shape[-1]264 265# Generate output266outputs = model.generate(**inputs, max_new_tokens=512)267response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)268 269# Parse output270processor.parse_response(response)271```272 273</details>274 275<details>276<summary>Code for processing Images</summary>277 278Make sure to install the following packages:279 280 281`pip install -U transformers torch torchvision accelerate`282 283You can then load the model with the code below:284 285```python286from transformers import AutoProcessor, AutoModelForMultimodalLM287 288MODEL_ID = "google/gemma-4-12B-it"289 290# Load model291processor = AutoProcessor.from_pretrained(MODEL_ID)292model = AutoModelForMultimodalLM.from_pretrained(293 MODEL_ID, 294 dtype="auto", 295 device_map="auto"296)297```298 299Once the model is loaded, you can start generating output by directly referencing the image URL in the prompt:300 301 302```python303# Prompt - add image before text304messages = [305 {306 "role": "user", "content": [307 {"type": "image", "url": "https://raw.githubusercontent.com/google-gemma/cookbook/refs/heads/main/apps/sample-data/GoldenGate.png"},308 {"type": "text", "text": "What is shown in this image?"}309 ]310 }311]312 313# Process input314inputs = processor.apply_chat_template(315 messages,316 tokenize=True,317 return_dict=True,318 return_tensors="pt",319 add_generation_prompt=True,320).to(model.device)321input_len = inputs["input_ids"].shape[-1]322 323# Generate output324outputs = model.generate(**inputs, max_new_tokens=512)325response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)326 327# Parse output328processor.parse_response(response)329```330 331</details>332 333 334<details>335<summary>Code for processing Videos</summary>336 337Make sure to install the following packages:338 339`pip install -U transformers torch torchvision librosa accelerate`340 341You can then load the model with the code below:342 343```python344from transformers import AutoProcessor, AutoModelForMultimodalLM345 346MODEL_ID = "google/gemma-4-12B-it"347 348# Load model349processor = AutoProcessor.from_pretrained(MODEL_ID)350model = AutoModelForMultimodalLM.from_pretrained(351 MODEL_ID, 352 dtype="auto", 353 device_map="auto"354)355```356 357Once the model is loaded, you can start generating output by directly referencing the video URL in the prompt:358 359 360```python361# Prompt - add video before text362messages = [363 {364 'role': 'user',365 'content': [366 {"type": "video", "video": "https://github.com/bebechien/gemma/raw/refs/heads/main/videos/ForBiggerBlazes.mp4"},367 {'type': 'text', 'text': 'Describe this video.'}368 ]369 }370]371 372# Process input373inputs = processor.apply_chat_template(374 messages,375 tokenize=True,376 return_dict=True,377 return_tensors="pt",378 add_generation_prompt=True,379).to(model.device)380input_len = inputs["input_ids"].shape[-1]381 382# Generate output383outputs = model.generate(**inputs, max_new_tokens=512)384response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)385 386# Parse output387processor.parse_response(response)388```389 390</details>391 392 393 394## **Best Practices**395 396For the best performance, use these configurations and best practices:397 398### 1. Sampling Parameters399 400Use the following standardized sampling configuration across all use cases:401 402* `temperature=1.0` 403* `top_p=0.95` 404* `top_k=64`405 406### 2. Thinking Mode Configuration407 408Compared to Gemma 3, the models use standard `system`, `assistant`, and `user` roles. To properly manage the thinking process, use the following control tokens:409 410* **Trigger Thinking:** Thinking is enabled by including the `<|think|>` token at the start of the system prompt. To disable thinking, remove the token. 411* **Standard Generation:** When thinking is enabled, the model will output its internal reasoning followed by the final answer using this structure: 412 `<|channel>thought\n`**[Internal reasoning]**`<channel|>` 413* **Disabled Thinking Behavior:** For all models except for the E2B and E4B variants, if thinking is disabled, the model will still generate the tags but with an empty thought block: 414 `<|channel>thought\n<channel|>`**[Final answer]**415 416> [!Note]417> Note that many libraries like Transformers and llama.cpp handle the complexities of the chat template for you.418 419### 3. Multi-Turn Conversations420 421* **No Thinking Content in History**: In multi-turn conversations, the historical model output should only include the final response. Thoughts from previous model turns must *not be added* before the next user turn begins.422 423### 4. Modality order424 425For optimal performance with multimodal inputs, place:426 427* Image content **before** the text in your prompt.428* Audio content **after** the text in your prompt.429 430### 5. Variable Image Resolution431 432Aside from variable aspect ratios, Gemma 4 supports variable image resolution through a configurable visual token budget, which controls how many tokens are used to represent an image. A higher token budget preserves more visual detail at the cost of additional compute, while a lower budget enables faster inference for tasks that don't require fine-grained understanding.433 434* The supported token budgets are: **70**, **140**, **280**, **560**, and **1120**. 435 * Use *lower budgets* for classification, captioning, or video understanding, where faster inference and processing many frames outweigh fine-grained detail. 436 * Use *higher budgets* for tasks like OCR, document parsing, or reading small text.437 438### 6. Audio439 440Use the following prompt structures for audio processing:441 442* **Audio Speech Recognition (ASR)**443 444```text445Transcribe the following speech segment in {LANGUAGE} into {LANGUAGE} text.446 447Follow these specific instructions for formatting the answer:448* Only output the transcription, with no newlines.449* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three.450```451 452* **Automatic Speech Translation (AST)**453 454```text455Transcribe the following speech segment in {SOURCE_LANGUAGE}, then translate it into {TARGET_LANGUAGE}.456When formatting the answer, first output the transcription in {SOURCE_LANGUAGE}, then one newline, then output the string '{TARGET_LANGUAGE}: ', then the translation in {TARGET_LANGUAGE}.457```458 459### 7. Audio and Video Length460 461All models support image inputs and can process videos as frames whereas the E2B, E4B, and 12B models also support audio inputs. Audio supports a maximum length of 30 seconds. Video supports a maximum of 60 seconds assuming the images are processed at one frame per second.462 463## **Model Data**464 465Data used for model training and how the data was processed.466 467### **Training Dataset**468 469Our pre-training dataset is a large-scale, diverse collection of data encompassing a wide range of domains and modalities, which includes web documents, code, images, audio, with a cutoff date of January 2025. Here are the key components:470 471* **Web Documents**: A diverse collection of web text ensures the model is exposed to a broad range of linguistic styles, topics, and vocabulary. The training dataset includes content in over 140 languages. 472* **Code**: Exposing the model to code helps it to learn the syntax and patterns of programming languages, which improves its ability to generate code and understand code-related questions. 473* **Mathematics**: Training on mathematical text helps the model learn logical reasoning, symbolic representation, and to address mathematical queries. 474* **Images**: A wide range of images enables the model to perform image analysis and visual data extraction tasks.475 476The combination of these diverse data sources is crucial for training a powerful multimodal model that can handle a wide variety of different tasks and data formats.477 478### **Data Preprocessing**479 480Here are the key data cleaning and filtering methods applied to the training data:481 482* **CSAM Filtering**: Rigorous CSAM (Child Sexual Abuse Material) filtering was applied at multiple stages in the data preparation process to ensure the exclusion of harmful and illegal content. 483* **Sensitive Data Filtering**: As part of making Gemma pre-trained models safe and reliable, automated techniques were used to filter out certain personal information and other sensitive data from training sets. 484* **Additional methods**: Filtering based on content quality and safety in line with [our policies](https://ai.google/static/documents/ai-responsibility-update-published-february-2025.pdf).485 486## **Ethics and Safety**487 488As open models become central to enterprise infrastructure, provenance and security are paramount. Developed by Google DeepMind, Gemma 4 undergoes the same rigorous safety evaluations as our proprietary Gemini models. 489 490### **Evaluation Approach**491 492Gemma 4 models were developed in partnership with internal safety and responsible AI teams. A range of automated as well as human evaluations were conducted to help improve model safety. These evaluations align with [Google’s AI principles](https://ai.google/principles/), as well as safety policies, which aim to prevent our generative AI models from generating harmful content, including:493 494* Content related to child sexual abuse material and exploitation 495* Dangerous content (e.g., promoting suicide, or instructing in activities that could cause real-world harm) 496* Sexually explicit content 497* Hate speech (e.g., dehumanizing members of protected groups) 498* Harassment (e.g., encouraging violence against people)499 500### **Evaluation Results**501 502For all areas of safety testing, we saw major improvements in all categories of content safety relative to previous Gemma models. Overall, Gemma 4 models significantly outperform Gemma 3 and 3n models in improving safety, while keeping unjustified refusals low. All testing was conducted without safety filters to evaluate the model capabilities and behaviors. For both text-to-text and image-to-text, and across all model sizes, the model produced minimal policy violations, and showed significant improvements over previous Gemma models' performance. 503 504## **Usage and Limitations**505 506These models have certain limitations that users should be aware of.507 508### **Intended Usage**509 510Multimodal models (capable of processing vision, language, and/or audio) have a wide range of applications across various industries and domains. The following list of potential uses is not comprehensive. The purpose of this list is to provide contextual information about the possible use-cases that the model creators considered as part of model training and development.511 512* **Content Creation and Communication** 513 * **Text Generation**: These models can be used to generate creative text formats such as poems, scripts, code, marketing copy, and email drafts. 514 * **Chatbots and Conversational AI**: Power conversational interfaces for customer service, virtual assistants, or interactive applications. 515 * **Text Summarization**: Generate concise summaries of a text corpus, research papers, or reports. 516 * **Image Data Extraction**: These models can be used to extract, interpret, and summarize visual data for text communications. 517 * **Audio Processing and Interaction**: The E2B, E4B, and 12B models can analyze and interpret audio inputs, enabling voice-driven interactions and transcriptions. 518* **Research and Education** 519 * **Natural Language Processing (NLP) and VLM Research**: These models can serve as a foundation for researchers to experiment with VLM and NLP techniques, develop algorithms, and contribute to the advancement of the field. 520 * **Language Learning Tools**: Support interactive language learning experiences, aiding in grammar correction or providing writing practice. 521 * **Knowledge Exploration**: Assist researchers in exploring large bodies of text by generating summaries or answering questions about specific topics.522 523### **Limitations**524 525* **Training Data** 526 * The quality and diversity of the training data significantly influence the model's capabilities. Biases or gaps in the training data can lead to limitations in the model's responses. 527 * The scope of the training dataset determines the subject areas the model can handle effectively. 528* **Context and Task Complexity** 529 * Models perform well on tasks that can be framed with clear prompts and instructions. Open-ended or highly complex tasks might be challenging. 530 * A model's performance can be influenced by the amount of context provided (longer context generally leads to better outputs, up to a certain point). 531* **Language Ambiguity and Nuance** 532 * Natural language is inherently complex. Models might struggle to grasp subtle nuances, sarcasm, or figurative language. 533* **Factual Accuracy** 534 * Models generate responses based on information they learned from their training datasets, but they are not knowledge bases. They may generate incorrect or outdated factual statements. 535* **Common Sense** 536 * Models rely on statistical patterns in language. They might lack the ability to apply common sense reasoning in certain situations.537 538### **Ethical Considerations and Risks**539 540The development of vision-language models (VLMs) raises several ethical concerns. In creating an open model, we have carefully considered the following:541 542* **Bias and Fairness** 543 * VLMs trained on large-scale, real-world text and image data can reflect socio-cultural biases embedded in the training material. Gemma 4 models underwent careful scrutiny, input data pre-processing, and post-training evaluations as reported in this card to help mitigate the risk of these biases. 544* **Misinformation and Misuse** 545 * VLMs can be misused to generate text that is false, misleading, or harmful. 546 * Guidelines are provided for responsible use with the model, see the [Responsible Generative AI Toolkit](https://ai.google.dev/responsible). 547* **Transparency and Accountability** 548 * This model card summarizes details on the models' architecture, capabilities, limitations, and evaluation processes. 549 * A responsibly developed open model offers the opportunity to share innovation by making VLM technology accessible to developers and researchers across the AI ecosystem.550 551**Risks identified and mitigations**:552 553* **Generation of harmful content**: Mechanisms and guidelines for content safety are essential. Developers are encouraged to exercise caution and implement appropriate content safety safeguards based on their specific product policies and application use cases. 554* **Misuse for malicious purposes**: Technical limitations and developer and end-user education can help mitigate against malicious applications of VLMs. Educational resources and reporting mechanisms for users to flag misuse are provided. 555* **Privacy violations**: Models were trained on data filtered for removal of certain personal information and other sensitive data. Developers are encouraged to adhere to privacy regulations with privacy-preserving techniques. 556* **Perpetuation of biases**: It's encouraged to perform continuous monitoring (using evaluation metrics, human review) and the exploration of de-biasing techniques during model training, fine-tuning, and other use cases.557 558### **Benefits**559 560At the time of release, this family of models provides high-performance open vision-language model implementations designed from the ground up for responsible AI development compared to similarly sized models.