CoolFace
Modelpublic

ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth

sourceHugging Facegemmaupdated 5mo agoView on Hugging Face
0likes28downloads
Model Card

Gemma4-E4B-Function-Calling-xLAM-Unsloth

This model is a fine-tuned version of Gemma4-E4B-it optimized for function calling using Unsloth for 2x faster training and 60% less VRAM.

Trained on the Salesforce/xlam-function-calling-60k dataset, which contains 60,000 function calling examples with queries, tool definitions, and structured answers.

Overview

PropertyValue
Developed byermiaazarkhalili
LicenseGEMMA
LanguageEnglish
Base ModelGemma4-E4B-it
Model Size4B parameters
Training FrameworkUnsloth + TRL
Training MethodSFT with QLoRA (4-bit)
Context Length2,048 tokens
GGUF AvailableGemma4-E4B-Function-Calling-xLAM-Unsloth-GGUF

Training Configuration

SFT + LoRA Settings

ParameterValue
Unsloth ClassFastModel
Chat Templategemma-4
Learning Rate2e-4
Batch Size2 per device
Gradient Accumulation4 steps
Effective Batch Size8
Max Steps1,000
OptimizerAdamW 8-bit
LR SchedulerLinear
Warmup Steps5
PrecisionAuto (BF16/FP16)
Gradient CheckpointingEnabled (Unsloth optimized)
Seed3407

LoRA Configuration

ParameterValue
LoRA Rank (r)16
LoRA Alpha16
LoRA Dropout0
Quantization4-bit QLoRA
Target Modulesattention + MLP (via FastModel)

Dataset

PropertyValue
DatasetxLAM Function Calling 60K
Training Samples60,000
FormatXML-tagged: <query>, <tools>, <answers>

Hardware

PropertyValue
GPUNVIDIA H100 80GB HBM3 (MIG 3g.40gb slice)
ClusterDRAC Fir (Compute Canada)
ExecutionPapermill on SLURM

Usage

Quick Start (Transformers)

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Check if the numbers 8 and 1233 are powers of two."}
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(response)

Using with Unsloth (Fastest)

python
from unsloth import FastModel

model, tokenizer = FastModel.from_pretrained(
    "ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth",
    max_seq_length=2048,
    load_in_4bit=True,
)

4-bit Quantized Inference

python
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch

quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
)

model = AutoModelForCausalLM.from_pretrained(
    "ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth",
    quantization_config=quantization_config,
    device_map="auto",
)

GGUF Versions

Quantized GGUF versions for CPU and edge inference are available at: [Gemma4-E4B-Function-Calling-xLAM-Unsloth-GGUF](https://huggingface.co/ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth-GGUF)

FormatDescription
Q4_K_MRecommended — good balance of quality and size
Q5_K_MHigher quality, slightly larger
Q8_0Near-lossless, largest GGUF size

Using with Ollama

bash
ollama pull hf.co/ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth-GGUF:Q4_K_M
ollama run hf.co/ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth-GGUF:Q4_K_M "Check if the numbers 8 and 1233 are powers of two."

Using with llama.cpp

bash
./llama-cli -m Gemma4-E4B-Function-Calling-xLAM-Unsloth-Q4_K_M.gguf -p "Check if the numbers 8 and 1233 are powers of two." -n 512

Limitations

  • —Language: Primarily trained on English data
  • —Knowledge Cutoff: Limited to base model's training data cutoff
  • —Hallucinations: May generate plausible-sounding but incorrect information
  • —Context Length: Fine-tuned with 2,048 token context window
  • —Safety: Not extensively safety-tuned; use with appropriate guardrails

Training Framework Versions

PackageVersion
Unsloth2026.4.4
TRL0.24.0
Transformers5.5.0
PyTorch2.9.0
Datasets4.3.0
PEFT0.18.1
BitsAndBytes0.49.2

Citation

bibtex
@misc{ermiaazarkhalili_gemma4_e4b_function_calling_xlam_unsloth,
    author = {ermiaazarkhalili},
    title = {Gemma4-E4B-Function-Calling-xLAM-Unsloth: Fine-tuned Gemma4-E4B-it with Unsloth},
    year = {2026},
    publisher = {Hugging Face},
    howpublished = {\url{https://huggingface.co/ermiaazarkhalili/Gemma4-E4B-Function-Calling-xLAM-Unsloth}}
}

Acknowledgments