tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF
Gemma-4 12B Coder — SFT v5 (GGUF)
gemma-4 12B coder for local, agentic tool use — GGUF quantizations for llama.cpp / Ollama.
Run it: llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M --jinja (full commands below).
⚠️ Tool-calling needs the recovery shim. The model emits gemma-4's native tool markup, whichllama.cpp --jinjaunder-parses — wrap your endpoint with the tool-shim (see Tool-calling below) to get standardtool_calls.
💡 Pick this for the best tool-calling (our gate winner). For an uncensored model, use SFT v5 + abliterated GGUF.
At a glance
Use it
# llama.cpp (server) — tool-calling needs the recovery shim, see below
llama-server -hf tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_M --jinja --ctx-size 16384
# Ollama
ollama run hf.co/tpls/gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-GGUF:Q4_K_MFiles
Sizes and a one-click loader are in the file browser / Quantizations widget above; the note says which quant to reach for.
Tool-calling
Tool-calling works — but llama.cpp --jinja doesn't recognise gemma-4's native tool-call markup, so the bare parser under-reports calls. The model is fine; the parser is blind to the format. Recover standard tool_calls with a small serve-side post-processor (no weight change, no latency beyond a regex scan).
Ready-to-use → [`tpls/gemma4-tool-shim`](https://huggingface.co/tpls/gemma4-tool-shim) — a drop-in callback for OpenAI-compatible proxies, a standalone (dependency-free) example, and the pure parser, all Apache-2.0, with the full recovery algorithm documented. Point your OpenAI-compatible endpoint through it.
You send tools the usual OpenAI way (tools=[…]); the model emits native markup; the shim turns it into a standard tool_calls object:
# model completion (raw):
<|tool_call>get_weather{"city": "Paris", "units": "celsius"}// after the shim:
{"finish_reason": "tool_calls",
"message": {"role": "assistant", "content": null,
"tool_calls": [{"id": "call_0", "type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\", \"units\": \"celsius\"}"}}]}}Tool-calling gate
Served the GGUF on llama.cpp (llama-server --jinja), prompted 7 tool-use cases + 1 no-tool abstain, scored whether a structured tool call was emitted. raw = llama.cpp native parse; shim = same outputs re-parsed for gemma-4 native markup. Tools folded into the prompt at eval time, matching training.
The rows are this model under two parse paths (raw and shim); the shim path is how it's served in production.
Intended use & limitations
Built for code generation and agentic tool use; serve locally via llama.cpp / Ollama, or use as a base to fine-tune / merge / quantize. Outputs can be wrong or fabricated — validate tool arguments before executing, and keep a human in the loop for anything consequential.
Where this sits in the family
- base (upstream) — `yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1`
- SFT v5 (weights)
- SFT v5 + abliterated (weights)
- SFT v5 + abliterated (GGUF)
- SFT v5 (GGUF) ← you are here
Provenance & reproduction
How this model was built — technique chain, training mix, and the exact knobs/pins, so the result is reproducible without any of our tooling.
Mechanics applied
1. sft-qlora
- tools_mode: mixed
2. imatrix-quant
- calibration: code + tool-call markup
- embed/output: kept at f16 (protects tool-call logits)
- eog_patch: tokens 105/106 → EOG (bounds the <|turn> runaway)
3. tool-shim
- where: a thin pre/post wrapper on the OpenAI-compatible endpoint
- format: re-parse
<|tool_call>NAME{json-args}(and leaked<|tool>…) intotool_calls
serve-side only — does not modify the weights; recommended for abliterated variants.
Training data & mix
Public sources; weights/row-caps are the exact balance.
Pinned revisions (byte-exact reproduction):
Agent-Ark/Toucan-1.5M(Kimi-K2) @0df3cf37f2abefb380370cfb02eabea2a35ae782Nanbeige/ToolMind(graphsyndatasets/graphsyn.jsonl) @8020ed1c03c367e4eb720ac3828ab4b0b95d8bafNousResearch/hermes-function-calling-v1(func-calling.json) @dae3e1d28cfbcf4b915c04ea1e072030529b4bdaNousResearch/hermes-function-calling-v1(func-calling-singleturn.json) @dae3e1d28cfbcf4b915c04ea1e072030529b4bdaSalesforce/xlam-function-calling-60k(toolsmode=mixed, toolsratio=0.5 (schemas folded into ~half the prompts)) @26d14ebfe18b1f7b524bd39b404b50af5dc97866
Training hyperparameters
Training environment
Exact pins the run trained against (the base arch needs a recent transformers).
Quantization environment
The GGUF bytes depend on the quantizer build, not just the weights — a different llama.cpp release rounds tensors differently and can change the convert mapping. Pins the toolchain these quants were produced with:
The image is the rolling:fulltag, not a digest — for byte-exact reproduction pin the image digest you build with. Theimatrix-quantstep above lists the calibration set and the EOG patch this build applied.
Part of the Gemma-4 12B Coder — active collection.
Something not right, or a request? Open a discussion — happy to help.
