CoolFace
Modelpublic

ajvikram/toolcall-2b-gguf

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
0likes348downloads
Model Card

Toolcall-2B — GGUF

Quantized builds of ajvikram/toolcall-2b, a 2B function-calling model fine-tuned from Qwen3.5-2B for local agent tool routing. Full results, training details and limitations are on the parent model's card.

FileSizeUse
toolcall-2b-Q4_K_M.gguf1.22 GBDefault. Smallest sensible quality loss, runs on a laptop CPU.
toolcall-2b-Q5_K_M.gguf1.35 GBA little closer to full precision for modest extra memory.
toolcall-2b-Q8_0.gguf1.93 GBNear-lossless; use when you have the memory.
toolcall-2b-f16.gguf3.63 GBUnquantized source for making your own quants.

Measured on the benchmark harness (safetensors, bf16): 36.35 overall on BFCL v4 against 33.85 for the Qwen3.5-2B base, with every group ahead of the base. The quantized builds are not separately scored.

Run it

bash
llama-server -m toolcall-2b-Q4_K_M.gguf --jinja -c 8192
bash
ollama run hf.co/ajvikram/toolcall-2b-gguf:Q4_K_M

The model uses Qwen3.5's native XML tool-call format, so any client that already parses Qwen3.5 tool calls works unchanged:

<tool_call>
<function=get_weather>
<parameter=city>
Berlin
</parameter>
</function>
</tool_call>

Verified with llama-cli on CPU: the Q4KM build loads, generates at roughly 33 tokens per second on an ARM CPU, and returns the call above for a get_weather tool given "What is the weather in Berlin?".

Thinking is off by default, matching how the model was trained and evaluated.

Notes

  • Built with llama.cpp (September 2026), which added Qwen3.5 conversion support; older builds cannot convert this architecture.
  • These are text-only builds. The base architecture is vision-capable, but this model was trained and evaluated purely on text tool calling.