CoolFace
Modelpublic

netarmy007/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes136downloads
Model Card

๐Ÿ’ป Gemma4-12B-Coder (GGUF) โ€” Composer 2.5 ร— Fable 5 โœจ

๐Ÿฃ Tiny footprint, big brain โ€” a local coding model for everyone

No matter your GPU. No matter your RAM. If you've got ~4.5 GB of VRAM or unified memory free, you can run your own private, offline coding assistant right now. ๐Ÿš€ This is the v1 / code edition โ€” distilled from real chain-of-thought so it thinks through a problem before writing the solution. ๐Ÿง ๐Ÿ’ป All local, all yours, no API, no cloud.

๐ŸŽฏ What it is

A focused fine-tune of Gemma 4 12B on verifiable Python coding data โ€” every training example's reasoning leads to code that actually passed its tests. The result reasons in the open (edge cases, complexity, approach) and then emits a clean, runnable solution. ๐Ÿ’š


๐Ÿ“Œ Announcements

๐Ÿ”œ safetensors (full-precision) upload coming. A lot of you want the safetensors master so you can roll your own quants / MLX builds โ€” it's on the way. My home Verizon connection keeps auto-dropping the upload partway through (it died on me again last night ๐Ÿ˜ค), so I'm heading to a Starbucks / library to push it through on a faster line. Thanks for the patience! ๐Ÿ™

๐Ÿš€ Big news โ€” v2 is almost here! Initial training of v2 is done, and it's now in benchmarking + final QA. I had Claude help me dig through a huge stack of the latest papers, and I'll be applying the newest methods to see how much further I can push it. ๐Ÿ“š So many of you flagged the agentic issues โ€” so this time I significantly increased the dataset (especially the agentic data). v2 is focused on agentic + coding. Stay tuned โ€” releasing this Friday or Saturday (US Pacific time)! ๐ŸŽ‰


๐Ÿ“ฃ Context length fixed: now 256K (was 131K) โ€” thanks, community! ๐Ÿ’š

A community member spotted that this model was reporting only a 131K context window. That turned out to be the well-known upstream Gemma 4 metadata bug โ€” Google's initial config.json shipped with max_position_embeddings: 131072 instead of the real 262144 (256K), and that value got baked into a lot of downstream finetunes and quants (including this one) before it was fixed upstream.

The weights were always fine โ€” it was purely a metadata field. All GGUF quants have been re-patched to the full 256K context (gemma4.context_length = 262144). Just re-download if you grabbed an earlier copy. ๐Ÿ™


๐Ÿ“š Training data (the interesting part ๐Ÿณ)

This is a distillation of two complementary chain-of-thought sources, both over verifiable Python coding tasks (algorithmic / function-level problems that come with deterministic tests):

  • โ€”*๐Ÿฅ‡ Main set โ€” Composer 2.5 real CoT. Genuine, model-authored reasoning traces. The teacher solved each problem, its code was run against the task's tests, and only the passing solutions were kept. So the reasoning you're learning from leads to code that actually works*.
  • โ€”๐Ÿฅˆ Aux set โ€” Fable 5 (released today! ๐ŸŽ‰). A clever twist: we took the problems where Composer 2.5 got it wrong and handed them to Fable 5 to redo โ€” re-deriving a fresh, self-consistent chain-of-thought and a correct solution, again gated on passing the tests. This recovers the hard cases the main teacher missed. These traces are synthetic (rationalized CoT), and are tagged separately so the two sources stay distinguishable.

The recipe: real CoT for the bulk of solid coverage, plus synthetic "second-attempt" CoT to patch the failures โ€” both verified by execution before anything entered training. โœ…


๐Ÿ“ฆ Pick your size (GGUF quants)

QuantSizeVibe
๐ŸŸข Q2_K4.5 GBtiniest โ€” runs almost anywhere
๐ŸŸก Q3_K_M5.7 GBgreat for 8 GB VRAM โ€” much better than Q2
๐Ÿ”ต Q4_K_M6.87 GBthe sweet spot ๐Ÿ‘Œ (recommended)
๐ŸŸฃ Q6_K9.11 GBnear-lossless
โšช Q8_011.8 GBbasically full quality

๐Ÿงฎ "Will it fit?" โ€” context length cheat-sheet

Rough estimates ๐Ÿค“ (assumes q8_0 KV cache + ~1.5 GB overhead; use `q4_0` KV cache for โ‰ˆ2ร— more context!). Max context is 256K. "โ€”" = won't fit, pick a smaller quant. โœ‚๏ธ

Your VRAM / unified mem๐ŸŸข Q2_K (4.5G)๐ŸŸก Q3_K_M (5.7G)๐Ÿ”ต Q4_K_M (6.87G)๐ŸŸฃ Q6_K (9.11G)โšช Q8_0 (11.8G)
8 GB~16K ctx~10Ktight (~2โ€“4K)โ€”โ€”
12 GB~48K~38K~30K~12Kโ€”
16 GB~80K~72K~64K~44K~22K
24 GB~200K~160K~128K~110K~88K
32 GB256K (max) ๐ŸŽ‰256K256K~230K~190K
๐Ÿ’ก Apple Silicon / integrated GPUs with unified memory count too โ€” same numbers, just slower than a dGPU. ๐Ÿ’ก Low on room? Drop a quant or switch KV cache to q4_0 and your context roughly doubles.

๐Ÿš€ How to run it (super easy)

Option A โ€” llama.cpp (recommended) ๐Ÿฆ™

  1. 1.Grab a quant above (e.g. โ€ฆ-Q4_K_M.gguf) and llama-server from llama.cpp.
โš ๏ธ Needs a recent llama.cpp (this is the gemma4_unified architecture โ€” older builds won't load it).
  1. 1.Run a server (Windows .bat shown โ€” tweak --port, --ctx-size to taste):
bat
@echo off
cd /d C:\llama.cpp
llama-server.exe ^
  -m C:\models\gemma4-coding-Q4_K_M.gguf ^
  --ctx-size 16384 ^
  --n-gpu-layers 99 ^
  --no-mmap ^
  -fa on ^
  --cache-type-k q8_0 --cache-type-v q8_0 ^
  --temp 1.0 --top-p 0.95 --top-k 64 ^
  --host 0.0.0.0 --port 18080
pause
  1. 1.Open http://localhost:18080 and chat. ๐ŸŽ‰ (Tip: bump --ctx-size per the table; use q4_0 KV for more.)

Option B โ€” one-click apps ๐Ÿ–ฑ๏ธ

Works in LM Studio, Jan, Ollama, etc. โ€” just import the GGUF, pick your quant, go. ๐Ÿพ

๐Ÿง  Thinking mode

This model thinks in Gemma's native thought channel before answering โ€” exactly how it was trained. Keep `enable_thinking=true` (the default chat template handles it). Recommended sampling: temp 1.0, top_p 0.95, top_k 64. For coding you can also go greedy (temp 0) for more deterministic solutions.


โš ๏ธ Good to know

  • โ€”Reduced refusals: the training data is task-focused with no safety hedging, so this refuses less than the base model. It is not safety-aligned โ€” add your own guardrails for production. Use responsibly. ๐Ÿ™
  • โ€”Specialized for Python / algorithmic coding. Reasoning quality is strongest in that domain; general-knowledge facts/numbers should still be double-checked.
  • โ€”English-centric.

๐Ÿ“š Base & License

  • โ€”License: Apache 2.0. Gemma 4 is released by Google under [Apache 2.0](https://ai.google.dev/gemma/apache_2) (unlike the older Gemma 1/2/3 terms), so this fine-tune is Apache 2.0 too โ€” free to use, modify, and redistribute. ๐ŸŽ‰
  • โ€”Base model: `google/gemma-4-12B-it`.
  • โ€”Personal/hobby project โ€” shared as-is, no warranty. Have fun, and happy hacking! ๐Ÿพโœจ