CoolFace
Modelpublic

iromu/Gemma3-1B-tools-GGUF

sourceHugging Facegemmaupdated 27d agoView on Hugging Face
0likes834downloads
Model Card

Gemma3 1B Tools GGUF

The Gemma 3 1B tool-calling model in GGUF format, fine-tuned with LoRA for tool calling and agent-style interactions.

Base model

This model was fine-tuned from:

google/gemma-3-1b-it

GGUF files

The model is provided in GGUF format at the following precisions:

PrecisionFile
BF16Gemma3-1B-tools-BF16.gguf (original precision)
Q4KMGemma3-1B-tools-Q4_K_M.gguf
Q5KMGemma3-1B-tools-Q5_K_M.gguf
Q8_0Gemma3-1B-tools-Q8_0.gguf

Training

Training was performed using NVIDIA NeMo AutoModel with LoRA/PEFT.

LoRA configuration

  • —LoRA dimension: 32
  • —LoRA alpha: 32
  • —Dropout: 0.05
  • —Target modules: *.proj (all *_proj linear layers)

Training configuration

  • —Max sequence length: 4096
  • —Learning rate: 5e-5 (cosine decay, 15 warmup steps, min 1e-6)
  • —Weight decay: 0.01
  • —Global batch size: 64 (micro batch 2 x 32 accumulation)
  • —Training steps: 336 (4 epochs)
  • —Mixed precision: bf16
  • —Validation loss: 0.579 → 0.4715 (final epoch)

Dataset

Training used the sft_tools split of the r0b0tlab/qwen3.8-max-glm5.2-kimi-k3-distillation dataset.

Tool-calling format

This model was trained with a custom chat template (embedded in the GGUF metadata). It renders the tool schemas into a developer turn and emits tool calls as:

<tool_call>
{"name": <function-name>, "arguments": <args-json-object>}
</tool_call>

Serving stacks must render prompts with this template for tool calling to work.

Intended use

  • —Structured tool/function calling
  • —Agent-style multi-step interactions
  • —Small-footprint on-device or edge deployment

It is not intended to be a general replacement for larger Gemma models.

GGUF versions

The model is available in GGUF format at:

  • —BF16
  • —Q4KM
  • —Q5KM
  • —Q8_0

Usage

Run the model with llama.cpp:

bash
llama-cli -hf iromu/Gemma3-1B-tools-GGUF:Q4_K_M

The BF16 GGUF file can be quantized locally to other GGUF precisions with llama-quantize if needed.

<!-- VALIDATION:BEGIN (auto-generated, do not edit) -->

Validation matrix

Tool-calling validation on the sft_tools validation split (greedy decoding, 384 max new tokens). Throughput is single-stream greedy decode, not serving throughput.

Pretrained base (google/gemma-3-1b-it): 2.0% exact-args match (1/50). Fine-tuned (BF16): 66.0% exact-args match (33/50) (+64pp vs base).

  • —GGUF-BF16: 20/50 (40.0%) exact, 66.0 tok/s — 61% of BF16.
  • —GGUF-Q4KM: 10/50 (20.0%) exact, 89.5 tok/s — 30% of BF16.
  • —GGUF-Q5KM: 24/50 (48.0%) exact, 60.7 tok/s — 73% of BF16.
  • —GGUF-Q8_0: 22/50 (44.0%) exact, 50.2 tok/s — 67% of BF16.
ModelQuantnTool call emittedNames matchExact args matchΔ exact vs BASEtok/s
Gemma3-1B-toolsBASE (google/gemma-3-1b-it)506/50 (12.0%)1/50 (2.0%)1/50 (2.0%)—68.5
Gemma3-1B-toolsBF165050/50 (100.0%)41/50 (82.0%)33/50 (66.0%)+64pp47.1
Gemma3-1B-toolsGGUF-BF165050/50 (100.0%)36/50 (72.0%)20/50 (40.0%)+38pp66.0
Gemma3-1B-toolsGGUF-Q4KM5050/50 (100.0%)19/50 (38.0%)10/50 (20.0%)+18pp89.5
Gemma3-1B-toolsGGUF-Q5KM5050/50 (100.0%)34/50 (68.0%)24/50 (48.0%)+46pp60.7
Gemma3-1B-toolsGGUF-Q8_05050/50 (100.0%)37/50 (74.0%)22/50 (44.0%)+42pp50.2

<!-- VALIDATION:END -->