CoolFace
Modelpublic

ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes81downloads
Model Card

Qwen3-8B-Function-Calling-xLAM-Unsloth

This model is a fine-tuned version of Qwen3-8B (Unsloth 4-bit) optimized for function calling using Unsloth for 2x faster training and 60% less VRAM.

Trained on the Salesforce/xlam-function-calling-60k dataset, which contains 60,000 function calling examples with queries, tool definitions, and structured answers.

Overview

PropertyValue
Developed byermiaazarkhalili
LicenseAPACHE-2.0
LanguageEnglish
Base ModelQwen3-8B (Unsloth 4-bit)
Model Size8B parameters
Training FrameworkUnsloth + TRL
Training MethodSFT with QLoRA (4-bit)
Context Length2,048 tokens
GGUF AvailableQwen3-8B-Function-Calling-xLAM-Unsloth-GGUF

Training Configuration

SFT + LoRA Settings

ParameterValue
Unsloth ClassFastLanguageModel
Chat Templatebuilt-in Qwen3
Learning Rate2e-4
Batch Size1 per device
Gradient Accumulation8 steps
Effective Batch Size8
Max Steps1 epoch (full dataset)
OptimizerAdamW 8-bit
LR SchedulerLinear
Warmup Steps5
PrecisionAuto (BF16/FP16)
Gradient CheckpointingEnabled (Unsloth optimized)
Seed3407

LoRA Configuration

ParameterValue
LoRA Rank (r)16
LoRA Alpha16
LoRA Dropout0
Quantization4-bit QLoRA
Target Modulesqproj, kproj, vproj, oproj, gateproj, upproj, down_proj

Dataset

PropertyValue
DatasetxLAM Function Calling 60K
Training Samples60,000
FormatXML-tagged: <query>, <tools>, <answers>

Hardware

PropertyValue
GPUNVIDIA H100 80GB HBM3 (MIG 3g.40gb slice)
ClusterDRAC Fir (Compute Canada)
ExecutionPapermill on SLURM

Training Outcome

MetricValue
SLURM Job ID36885898
Runtime3h 48m 36s (13716s)
Final Training Loss0.2186
Peak VRAM17.07 GB
GPUH100 80GB HBM3 (MIG 3g.40gb)

Usage

Quick Start (Transformers)

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Check if the numbers 8 and 1233 are powers of two."}
]

text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
response = tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(response)

Using with Unsloth (Fastest)

python
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    "ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth",
    max_seq_length=2048,
    load_in_4bit=True,
)

4-bit Quantized Inference

python
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch

quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
)

model = AutoModelForCausalLM.from_pretrained(
    "ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth",
    quantization_config=quantization_config,
    device_map="auto",
)

GGUF Versions

Quantized GGUF versions for CPU and edge inference are available at: [Qwen3-8B-Function-Calling-xLAM-Unsloth-GGUF](https://huggingface.co/ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth-GGUF)

FormatDescription
Q4_K_MRecommended — good balance of quality and size
Q5_K_MHigher quality, slightly larger
Q8_0Near-lossless, largest GGUF size

Using with Ollama

bash
ollama pull hf.co/ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth-GGUF:Q4_K_M
ollama run hf.co/ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth-GGUF:Q4_K_M "Check if the numbers 8 and 1233 are powers of two."

Using with llama.cpp

bash
./llama-cli -m Qwen3-8B-Function-Calling-xLAM-Unsloth-Q4_K_M.gguf -p "Check if the numbers 8 and 1233 are powers of two." -n 512

Limitations

  • —Language: Primarily trained on English data
  • —Knowledge Cutoff: Limited to base model's training data cutoff
  • —Hallucinations: May generate plausible-sounding but incorrect information
  • —Context Length: Fine-tuned with 2,048 token context window
  • —Safety: Not extensively safety-tuned; use with appropriate guardrails

Training Framework Versions

PackageVersion
Unsloth2026.4.4
TRL0.24.0
Transformers5.5.0
PyTorch2.9.0
Datasets4.3.0
PEFT0.18.1
BitsAndBytes0.49.2

Citation

bibtex
@misc{ermiaazarkhalili_qwen3_8b_function_calling_xlam_unsloth,
    author = {ermiaazarkhalili},
    title = {Qwen3-8B-Function-Calling-xLAM-Unsloth: Fine-tuned Qwen3-8B (Unsloth 4-bit) with Unsloth},
    year = {2026},
    publisher = {Hugging Face},
    howpublished = {\url{https://huggingface.co/ermiaazarkhalili/Qwen3-8B-Function-Calling-xLAM-Unsloth}}
}

Acknowledgments