CoolFace
Modelpublic

twinkle-ai/gemma-3-4B-T1-it-GGUF

sourceHugging Facegemmaupdated 8mo agoView on Hugging Face
4likes285downloads
Model Card

Gemma 3 4B T1-it GGUF Collection

<div align="center" style="line-height: 1;"> <a href="https://discord.gg/Cx737yw4ed" target="blank" style="margin: 2px;"> <img alt="Discord" src="https://img.shields.io/badge/Discord-Twinkle%20AI-7289da?logo=discord&logoColor=white&color=7289da" style="display: inline-block; vertical-align: middle;"/> </a> <a href="https://huggingface.co/twinkle-ai" target="blank" style="margin: 2px;"> <img alt="Hugging Face" src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Twinkle%20AI-ffc107?color=ffc107&logoColor=white" style="display: inline-block; vertical-align: middle;"/> </a> <!-- Gemma 模型在 Hugging Face 上為 gated,使用者需同意 Google usage license --> <a href="https://huggingface.co/google/gemma-3-4b-pt" style="margin: 2px;"> <img alt="License" src="https://img.shields.io/badge/License-gemma-f5de53?&color=0081fb" style="display: inline-block; vertical-align: middle;"/> </a> </div>

GGUF quantized models converted from twinkle-ai/gemma-3-4B-T1-it for use with llama.cpp.

Gemma3-4B-T1-it

About

Gemma 3 4B T1-it is a small language model fine-tuned on Taiwan-focused datasets, supporting both English and Traditional Chinese. This repository provides multiple quantization formats optimized for different use cases.

Available Models

ModelSizeUse Case
twinkle-ai-gemma-3-4B-T1-it-BF16.ggufLargestBest quality, highest precision
twinkle-ai-gemma-3-4B-T1-it-F16.ggufLargeHigh quality, good precision
twinkle-ai-gemma-3-4B-T1-it-Q8_0.ggufMediumBalanced quality and speed
twinkle-ai-gemma-3-4b-t1-it-q4_k_m.ggufSmallestFastest inference, lower memory

Quick Start

Option 1: Using Hugging Face Hub (Recommended)

Install llama.cpp via Homebrew:

bash
brew install llama.cpp

Run inference directly from Hugging Face:

bash
llama-cli --hf-repo thliang01/gemma-3-4B-T1-it-Q8_0-GGUF \
  --hf-file gemma-3-4b-t1-it-q8_0.gguf \
  -p "Your prompt here"

Start as a server:

bash
llama-server --hf-repo thliang01/gemma-3-4B-T1-it-Q8_0-GGUF \
  --hf-file gemma-3-4b-t1-it-q8_0.gguf \
  -c 2048

Option 2: Build from Source

Step 1: Clone llama.cpp repository
bash
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
Step 2: Build llama.cpp

Basic build (CPU only):

bash
LLAMA_CURL=1 make

Hardware-specific build options:

  • —NVIDIA GPU (Linux):
bash
  LLAMA_CUDA=1 LLAMA_CURL=1 make
  • —Apple Silicon (Mac):
bash
  LLAMA_METAL=1 LLAMA_CURL=1 make
  • —AMD GPU (ROCm):
bash
  LLAMA_HIPBLAS=1 LLAMA_CURL=1 make
Step 3: Run inference
bash
./llama-cli --hf-repo thliang01/gemma-3-4B-T1-it-Q8_0-GGUF \
  --hf-file gemma-3-4b-t1-it-q8_0.gguf \
  -p "Your prompt here"
Step 4: Start server (optional)
bash
./llama-server --hf-repo thliang01/gemma-3-4B-T1-it-Q8_0-GGUF \
  --hf-file gemma-3-4b-t1-it-q8_0.gguf \
  -c 2048

Advanced Usage

Choosing the Right Model

Select a model based on your needs:

  • —Best Quality: Use BF16 or F16 versions (requires more memory)
  • —Balanced: Use Q8_0 version (recommended for most users)
  • —Resource Constrained: Use q4_k_m version (suitable for devices with limited memory)

Common Parameters

  • —-p "prompt": Your input text for the model to respond to
  • —-c 2048: Context length (maximum number of tokens that can be processed)
  • —--hf-repo: Hugging Face repository name
  • —--hf-file: Model file name to use

Adjusting Generation Parameters

bash
llama-cli --hf-repo thliang01/gemma-3-4B-T1-it-Q8_0-GGUF \
  --hf-file gemma-3-4b-t1-it-q8_0.gguf \
  -p "Your prompt here" \
  --temp 0.7 \
  --top-p 0.9 \
  --repeat-penalty 1.1

Parameter explanations:

  • —--temp: Temperature (0.0-2.0), higher values produce more random output
  • —--top-p: Nucleus sampling parameter (0.0-1.0)
  • —--repeat-penalty: Repetition penalty to avoid repetitive content

Model Information

  • —Base Model: twinkle-ai/gemma-3-4B-T1-it
  • —Languages: English, Traditional Chinese
  • —License: Gemma
  • —Format: GGUF (converted via GGUF-my-repo)

Training Data

  • —Taiwan reasoning and instruction datasets
  • —Contract review and legal documents
  • —Multimodal and long-form content
  • —Instruction-following examples

Benchmarks

  • —TMMLU+: 47.44% accuracy
  • —MMLU: 59.13% accuracy
  • —TW Legal Benchmark: 44.18% accuracy

Troubleshooting

Common Issues

Q: Getting out of memory errors?

A: Try using a smaller quantized version like q4_k_m, or reduce the context length parameter -c.

Q: How can I speed up inference?

A:

  1. 1.Use GPU acceleration (add hardware-specific flags during compilation)
  2. 2.Choose a smaller quantized model (like q4_k_m)
  3. 3.Reduce context length

Q: What prompt format does the model support?

A: This is an instruction-tuned model. Use a clear instruction format, for example:

text
Please analyze the main clauses of the following contract: [contract content]

Links

Contributing

If you have any questions or suggestions, please feel free to open a discussion in the Hugging Face repository.


Note: On first run, llama.cpp will automatically download the model file from Hugging Face. Please ensure you have a stable internet connection.