CoolFace
Modelpublic

DuoNeural/translategemma-4b-it-LiteRT

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes77downloads
Model Card

language:

  • —en tags:
  • —duoneural
  • —litert
  • —edge
  • —gguf
  • —on-device
  • —translategemma
  • —gemma
  • —translation
  • —multilingual
  • —litert
  • —edge basemodel: google/translategemma-4b-it pipelinetag: text-generation license: other ---

# translategemma-4b-it-LiteRT

TranslateGemma 4B Instruct — multilingual translation for on-device inference — converted for mobile and edge deployment by DuoNeural.

  • —Source model: google/translategemma-4b-it
  • —Format: GGUF Q4KM (llama.cpp-compatible)
  • —File size: 2490 MB
  • —Quantization: 4-bit K-mean (Q4KM) — excellent accuracy/size trade-off for edge devices
  • —Target platforms: Android, iOS, desktop edge inference
  • —Converted: 2026-05-06 06:06:48 by Archon / DuoNeural

## Usage

### llama.cpp (CLI)

bash
    ./llama-cli -m translategemma-4b-it-LiteRT_Q4_K_M.gguf -n 512 --temp 0.7

### Google AI Edge / MediaPipe (Android/iOS) This GGUF is compatible with MLC-LLM and llama.cpp Android bindings for on-device inference. For use with Google Edge Gallery, convert to .task bundle using MediaPipe LLM conversion tools.

### Python via llama-cpp-python

python
    from llama_cpp import Llama

    llm = Llama(
        model_path="translategemma-4b-it-LiteRT_Q4_K_M.gguf",
        n_ctx=2048,
        n_threads=4,
        verbose=False,
    )

    response = llm.create_chat_completion(
        messages=[
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": "Hello! How can you help me today?"},
        ]
    )
    print(response["choices"][0]["message"]["content"])

### Ollama

bash
    ollama run hf.co/DuoNeural/translategemma-4b-it-LiteRT

## About the Conversion

Converted using llama.cpp GGUF pipeline with CUDA acceleration. Source weights downloaded from HuggingFace, converted to F16 GGUF, then quantized to Q4KM.


## DuoNeural

DuoNeural is an open AI research lab — human + AI in collaboration.

PlatformLink
HuggingFacehuggingface.co/DuoNeural
Websiteduoneural.com
GitHubgithub.com/DuoNeural
X / Twitter@DuoNeural
Emailduoneural@proton.me
Newsletterduoneural.beehiiv.com
Supportbuymeacoffee.com/duoneural

### DuoNeural Research Publications

Open access, CC BY 4.0. Authored by Archon, Jesse Caldwell, Aura — DuoNeural.

### Research Team

  • —Jesse — Vision, hardware, direction
  • —Archon — Lab Director, post-training, abliteration, experiments
  • —Aura — Research AI, literature synthesis, novel proposals

Subscribe to the lab newsletter at [duoneural.beehiiv.com](https://duoneural.beehiiv.com) for model drops before they go anywhere else.