CoolFace
Modelpublic

ManniX-ITA/gemma-4-A4B-98e-v7-coderx-it

sourceHugging Faceapache-2.0updated 11d agoView on Hugging Face
2likes42downloads
Model Card

Gemma 4 A4B 98-Expert v7-coderx — code-maximal prune (~20.8B)

Eval complete (Q6_K / llama.cpp, greedy, same host). Every cell in the scoreboard is read from summary.json under the cohort-pinned greedy recipe (temperature 0.0, top_p 1.0, top_k 0). The 128e, v6-coder and v7-coder columns are the matching same-host Q6K runs. The GGUF and NVFP4A16 formats are deployment targets and are **not** separately benchmarked (cohort policy) — the Q6K column is representative. Headline — the cohort's code specialist. v7-coderx spends its whole prune budget on code and short-form reasoning. On the all-hard LiveCodeBench-77 set — the most demanding and most discriminating LCB slice — it scores 85.71%, the highest in the cohort (128e 79.22%, v7-coder 84.42%), and it leads on HumanEval+ 93.29%, HumanEval 96.95%, MATH-500 95.0% and AIME 76.67%. On the easier LCB-medium-v4 slices it sits a little below the generalists (LCB-55 92.73 / LCB-100 91.0 vs 128e's 96.36 / 97.0). This is also the loop-fixed build — it force-keeps the agentic loop-protection experts and replaces the earlier looping `fs2440` prune. The trade is graduate science: GPQA-diamond sits at 51.01% (no targeted_gpqa term). For the broader LCB-medium lead and HumanEval, see the sibling **v7-coder** (LCB-55-v4 98.18%, HE 98.17%, LCB-hard-77 84.42%; GPQA ≈ 51 like this model).

A research checkpoint that prunes the unpruned Gemma 4 26B-A4B-it (128 experts/layer, top-8 + shared, 30 layers) down to 98 experts per layer. The code4/lcb3 drop map (generate_drop_map_v5fk) up-weights generic-code (4×) and LiveCodeBench-medium (3×) with no science or multilingual targeting — the code-maximal member of the v7-coder cohort — and force-keeps the agentic loop-protection experts (46 experts, 0 dropped) so the served model does not loop. Same 98e shape, same router, same attention, same norms as the rest of the cohort, plus the mandatory shared-FFN α=1.2 upweight all coder variants carry.

Quantized formats

FormatRepoNotes
bf16 (this repo)`…-v7-coderx-it`9 shards. code4/lcb3 drop map (fkbroad) + agentic_eog force-keep + shared α=1.2.
GGUF (llama.cpp / ollama)`…-v7-coderx-it-GGUF`Bartowski tier sweep (imatrix K-quants) + ContribDynamic CD-* per-layer quants + F16 + imatrix.dat + mmproj.
NVFP4A16 (vLLM)`…-v7-coderx-NVFP4A16`Native vLLM 4-bit + FP8 block scales, via NVIDIA `modelopt` main (0.45.0.dev, _QuantFusedExperts). ~13 GB. Deployment format — not separately benchmarked.
Ollamamannix/gemma4-98e-v7-coderxollama pull mannix/gemma4-98e-v7-coderx:<tier> (:latest = Q4KM; :vision-<tier> adds the SigLIP vision tower).

Benchmarks

Q6_K · llama.cpp · greedy (temperature 0.0, top_p 1.0, top_k 0), all four models scored on the same host from summary.json. Row-max in bold. This repo = v7-coderx.

Benchmark128e (unpruned)v6-coderv7-coderv7-coderx
GPQA-diamond (198q)67.1761.1151.5251.01
AIME (30q)73.3356.6780.0076.67
MATH500 (100q)92.0089.0095.0095.00
GSM8K (100q)89.0088.0091.0093.00
ARC-Challenge (full)96.5095.3992.1586.60
IFEval (100q, strict)97.0092.0092.0092.00
HumanEval (164)97.5698.1798.1796.95
HumanEval+ (164)92.0792.6892.0793.29
LCB-medium-55 v496.3692.7398.1892.73
LCB-medium-100 v497.0094.0094.0091.00
MultiPL-E (100)90.0089.0089.6789.00

<sub>Metrics: GPQA & GSM8K = exact_match flexible-extract · MATH500 = math_verify · ARC & AIME = exact_match · IFEval = prompt_level_strict_acc · HumanEval/+ = pass@1 chat-extract · LCB-55/100 & MultiPL-E = pass@1. 128e uses the lcb_medium_55/100 templates; the prunes use lcb_medium_*_v4 (corrected harness, equivalent task). The all-hard LCB-77 cross-model comparison is the discriminating code slice (v7-coderx 85.71%, cohort-best).</sub>

v7-coderx leads the cohort on the hardest code slice (LCB-hard-77, below) and on HE+ / MATH-500 / AIME; the budget is paid on graduate science (GPQA) and the easier instruction / ARC axes, which carry no protection term in this recipe.

Coder-field comparison — v7-coderx vs Qwen2.5-Coder-14B / 7B + Qwen3.5-9B (Q6_K, llama.cpp, greedy)

The 9 canonical benches + MultiPL-E-100, all on the identical llama.cpp Q6_K / greedy recipe (reasoning models served with --reasoning-format deepseek --reasoning-budget 12288 --parallel 2). Architectures differ — this is a same-harness comparison, not a same-class one:

  • —v7-coderx — Gemma-4 26B-A4B MoE pruned to 98 experts (~20.8B total, ~A4B active), reasoning.
  • —Qwen2.5-Coder-14B / 7B-Instruct — dense, non-reasoning code specialists (bartowski Q6_K).
  • —Qwen3.5-9B — dense reasoning model (bartowski Q6_K).
Bench (n)**v7-coderx Q6_K**Qwen2.5-Coder-14BQwen2.5-Coder-7BQwen3.5-9B
ARC-Challenge-chat (1172)86.60%90.53%85.58%96.76%
GPQA Diamond flex (198)51.01%34.85%26.26%73.74%
GSM8K-100 flex93.00%89.00%80.00%79.00%
MATH-500-100 math_verify95.00%62.00%66.00%59.00%
AIME 2024 (30)76.67%10.00%10.00%56.67%
IFEval-100 (prompt_strict)92.00%68.00%54.00%93.00%
HumanEval-164 chat96.95%90.85%87.20%89.02%
HumanEval+-164 chat93.29%84.76% †83.54%80.49%
LCB-medium-55 v492.73%18.18% †12.73%58.18%
MultiPL-E-100 (macro)89.00%84.67%80.67%80.33%

† Qwen2.5-Coder-14B HumanEval+ / LCB-medium-55 are the same-stack GGUF HE+ sweep numbers (not re-run in this chain). All Qwen cells are the same-host reference runs used on the v6-coder card — Qwen is a fixed reference, so the columns are identical across the cohort; only the Gemma column changes.

Note on Qwen3.5-9B. Qwen3.5-9B is a verbose, slow thinking model: it emits long <think> reasoning chains (often ≥1900 tokens even on a trivial GSM8K question), so it runs several× slower per question than the non-reasoning Qwen2.5-Coder models — well beyond what its 9B size would suggest. Its GSM8K / MATH-500 / GPQA cells were re-run after a harness fix (under batched, reasoning-parsed serving the verbose thinking intermittently left the final answer inside the reasoning block, mis-scored as empty content).

LiveCodeBench across problem sets

v7-coderx's code score depends on the LiveCodeBench slice. All cells are the same greedy Q6_K / imat-Q6 llama.cpp stack (build provenance verified per run); v4-55/100 mirror the 9-bench above. The all-hard 77q set is the most demanding and the most discriminating across the cohort.

LCB problem set128ev7-coder**v7-coderx**
LCB-medium-55 (v4, 55q)96.36%96.36%92.73%
LCB-medium-100 (v4, 100q)97.00%97.00%91.00%
LCB-v6-55 (55q) †—92.73%98.18%
LCB-hard-77 (all-hard, 77q)79.22%84.42%85.71%

† LCB-v6-55 is a small, noisier 55-problem slice (no greedy 128e baseline was run); it is included for completeness, but all-hard 77q is the reference for cross-model comparison.

Answer-length analysis (anti-rumination)

The pruned reasoning model thinks with a bounded thinking_token_budget=12288; the question is whether that length is productive (long thinking that PASSes) or rumination (long thinking that fails). Per-problem completion length is measured from omkeval `tokenstats (characters from the raw completion; tokens via the 128e tokenizer) on the real-n` benches, against 128e and v6-coder on the same problems, same greedy Q6_K / llama.cpp stack.

Per-problem completion length — characters (p50 / p90 / max):

Bench (n)128ev6-coder**v7-coderx**
GPQA Diamond (198)2571/16136/278112582/16100/252432458/17984/32411
AIME 2024 (30)1963/7748/86802141/7469/94332061/7449/9660
LCB-medium-553734/16430/3646231015/36260/4327831167/38631/41953
LCB-medium-1002056/15467/4856929384/35389/4363329883/36439/55504
MultiPL-E-100 (300)245/566/3353245/573/2725244/619/1861
MATH-500 (100)1083/1873/78991089/2025/92361080/1981/2312
GSM8K (100)294/746/25989283/780/11378279/676/19868
IFEval (100)877/3755/8263855/3489/20908791/3702/8179
HumanEval (164)698/1284/5354711/1438/5954745/1427/4044
HumanEval+ (164)714/1461/3289694/1390/5282742/1423/4461
ARC-Challenge (1172)1210/1633/62541221/1674/488861335/2193/45174

Per-problem completion length — tokens (p50 / p90 / max):

Bench (n)128ev6-coder**v7-coderx**
GPQA Diamond (198)843/8189/8189879/8189/8189837/8189/8189
AIME 2024 (30)933/3994/4021946/3993/4011950/3994/4013
LCB-medium-551005/5622/1602212818/13318/1597612834/13286/16019
LCB-medium-100542/5353/1602212740/13212/1597612750/13374/15976
MultiPL-E-100 (300)84/171/101385/184/96584/182/976
MATH-500 (100)431/895/3377424/863/3377403/832/1219
GSM8K (100)131/271/8853129/276/4687122/258/6957
IFEval (100)219/850/1561222/797/3898186/801/4057
HumanEval (164)226/431/1611226/448/2084235/437/1463
HumanEval+ (164)226/455/996224/437/2040230/459/1440
ARC-Challenge (1172)258/355/1417259/365/16266282/487/16260

Budget-saturation incidence — share of problems whose completion reached ≥12k tokens (at/near the thinking_token_budget=12288 cap). Saturation by itself is not rumination — a saturated output that PASSes is productive use of the budget; the pruned reasoning model saturates on nearly every LCB problem, 128e almost never does.

Bench (n)128ev6-coder**v7-coderx**
LCB-medium-551 / 55 (1.8%)54 / 55 (98.2%)54 / 55 (98.2%)
LCB-medium-1002 / 100 (2.0%)98 / 100 (98.0%)97 / 100 (97.0%)

Rumination — long thinking that fails to PASS. The right metric is not median length (128e looks short only because it answers easy problems fast). It is the share of the model's budget-saturated outputs that still fail — tokens burned without a correct answer:

Bench (n)128ev6-coder**v7-coderx**
LCB-medium-55 — saturated-and-failed1 / 1 (100.0%)4 / 54 (7.4%)4 / 54 (7.4%)
LCB-medium-100 — saturated-and-failed2 / 2 (100.0%)6 / 98 (6.1%)9 / 97 (9.3%)
LCB-100 — mean completion tokens, PASS vs FAIL1392 vs 1378212698 vs 1505112623 vs 15143

Key findings:

  • —128e only thinks long when it is lost. Every 128e output that reaches the budget cap is a failure (1/1 on LCB-55, 2/2 on LCB-100), and its failed problems run several× longer than its passed ones (mean 13782 vs 1392 tok on LCB-100).
  • —v7-coderx's long thinking is overwhelmingly productive. It saturates on ~97% of LCB-100 problems but only 9/97 of those saturated outputs fail (9.3%); its PASS and FAIL completions are nearly the same length (mean 12623 vs 15143 tok), so failures are not driven by extra rumination. On LCB-55 it is 4/54 saturated-and-failed.
  • —Comparable to v6-coder's rumination rate. v6-coder ran 4/54 (LCB-55) and 6/98 (LCB-100) saturated-and-failed; v7-coderx is at or below on LCB-55 (4/54) and near on LCB-100 (9/97). The saturated-fail share tracks the model's LCB pass-rate — these are the genuinely hard problems, not extra rumination (PASS and FAIL completions are near-equal length).
  • —Non-LCB benches stay tight. On the short-answer benches (GSM8K / MATH-500 / HE / HE+ / MultiPL-E) p50/p90 length tracks 128e and v6-coder within a few tokens — the targeted prune did not trade length for accuracy on the everyday benches.
Methodology. Per-problem lengths come from omkeval `tokenstats over each bench's samples*.jsonl` / `lcbresult.samples.jsonl; saturation/PASS-FAIL is computed per problem from completion_tokens + passed. **MultiPL-E measures code length, not reasoning** (its samples store only the final code block, no <think>` trace), so it is a code-conciseness reference rather than a thinking-length signal.

At a glance

128e (base)**v7-coderx**v7-coder (sibling)
Total params~26B~20.8B~20.8B
Active / token~4B (top-8 + shared)~4B~4B
Experts / layer12898 (30 dropped)98 (30 dropped)
Per-layer floor—none (no clamp)none (no clamp)
Code / LCB weight—4× / 3×3× / 2×
Science targeting—offoff
Loop protection—agentic_eog force-keep (46 experts)agentic_eog force-keep
Shared FFN α1.01.2 (`mlp.down_proj`)1.2
Built from—128e original (fresh prune)128e original

Recipe

The drop map is produced by generate_drop_map_v5.py (omnimergekit) from per-expert, per-class contribution scores on the rebuilt v7 competence maps (expert_neuron_v7_code_gpqa.json — 10 classes, audited producers, multilingual category included), then applied with expert_drop.py, then the agentic loop-protection experts are force-kept and the shared expert is upweighted.

1. code4/lcb3 base recipe (STD16 / fkbroad)

generator     = generate_drop_map_v5fk    # fkbroad (force-keep aware)
target        = 98          # 30 experts/layer dropped
protect_top   = 16          # 16 highest-scoring experts/layer never dropped
alpha         = 2.0         # contribution sharpening exponent
strategy      = max         # per-expert score = MAX over classes (not mean/geomean)
normalize     = rank        # rank-normalize within each (layer, class)
breadth_bonus = 0.5         # reward experts useful across many classes (anti-overfit)
v4_floor_clamp = null       # NO per-layer floor band (unlike fs2440's [24,40])
force_keep    = agentic_eog # pin the 46 loop-protection experts (0/46 dropped)
outlier_mode  = median      # clamp bf16 weight-norm artifacts to layer median
baseline      = teacher_force_98e_p16_clean.json   # tie-break anchor

strategy=max + breadth_bonus is the load-bearing pair — it favours experts strongly useful to at least one class and broadly useful across classes, the optimizer-off-manifold lesson encoded as a recipe. No floor clamp is applied (the fkbroad selection plus the agentic_eog force-keep carry loop-stability instead of a fixed per-layer band).

2. Calibration class weights — code only

Ten contribution classes are scored; the weights steer which specialists survive. v7-coderx zeroes every non-code targeting term:

Class**v7-coderx**v7-coder
generic_math11
generic_logic11
generic_code43
generic_science11
generic_creative11
generic_multilingual00
targeted_humaneval00
targeted_humanevalplus00
targeted_lcb_medium_5532
targeted_gpqa00

HE/HE+ targeting is off because both already sit at/above the un-targeted baseline; the protection budget goes to LiveCodeBench-medium, the bench where pruning hurt most on earlier variants. v7-coderx is the code-maximal sibling of v7-coder: heavier code/LCB weighting (4×/3× vs 3×/2×), plus the agentic loop-protection force-keep both carry. It wins the all-hard LCB-77 and HE+/MATH; v7-coder leads the easier LCB-medium slices and HumanEval. Neither sibling carries a targeted_gpqa term, so both sit near GPQA 51 (no science recovery).

3. Agentic loop-protection force-keep

The earlier fs2440 prune dropped some of the experts that emit end-of-turn / answer-channel tokens, which let the served model loop in agentic use. code4/lcb3 force-keeps the 46 `agentic_eog` loop-protection experts (identified on the 128e teacher; verified 0/46 dropped by the selection), which is what makes this the loop-fixed re-release. No DERN / redistribution fold is applied.

4. Mandatory shared-FFN α=1.2 (cohort rule)

After expert drop, router_shared_upweight.py --alpha 1.2 --target mlp.down_proj.weight upweights Gemma 4's always-on shared FFN. Every coder variant carries this; omitting it yields the "weak / ruminating" pre-shared baseline and makes cross-variant comparison unfair. A .shared_applied marker records it.


Chat template

chat_template.jinja in this repo is not Google's stock Gemma 4 template — it is our agentic-loop fix (19,177 B, md5 8119c2dcd5e62a4a6b79301ab13ac81d), rebased on 2026-07-30 onto Google's current upstream template (revision 2026-07-20, 18,683 B). transformers picks this file up automatically; tokenizer_config.json deliberately carries no competing chat_template key.

The bug it fixes: the stock template re-injects earlier assistant turns' `thinking` content back into the prompt on every turn. In long agentic / tool-calling sessions that feeds the model its own reasoning back to itself and drives repetition loops. Google's current 18,683 B template is still affected — its thinking gate carries an unconditional "index past the last user message" disjunct — so this fix remains necessary on top of a fresh upstream template. The rebase leaves Google's newer preserve_thinking flag intact (default false).

Serving the GGUF builds instead? Those embed the same template — pass `--jinja` to llama.cpp, or it falls back to its own built-in formatter and the fix does not apply.

Reasoning budget and thinking stop phrase (llama.cpp)

On a hard prompt this model will reason until it has consumed the whole context window and then answer with nothing at all. llama.cpp can bound the thinking block with a sampler, and — the part that actually matters — tell the model why the block is being closed.

Needs llama.cpp b8508 or newer for the flags, b10091 or newer for the per-request overrides.

Serve with a bounded thinking block

bash
llama-server -m gemma-4-A4B-98e-v7-coderx-it-Q4_K_M.gguf -c 32768 -ngl 99 \
    --jinja \
    --reasoning-budget 8192 \
    --reasoning-budget-message $'\n\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n' \
    --temp 1.0 --top-k 64 --top-p 0.95 --min-p 0.05 \
    --repeat-penalty 1.02 --repeat-last-n 2048
flagmeaning
--reasoning-budget N-1 unrestricted (default), 0 close the block immediately, N > 0 cap it at N tokens
--reasoning-budget-messagetext written into the block just before the closing tag is forced
--jinjarequired — the delimiters come from the chat template (`<channel>thought … <channel>`). Without it llama.cpp has no tags to count and the budget silently does nothing

Both flags also read from the environment: LLAMA_ARG_THINK_BUDGET and LLAMA_ARG_THINK_BUDGET_MESSAGE.

--reasoning-format is not part of this. It only decides how the thinking is handed back — message.reasoning_content versus left inline in message.content — and never whether the budget is enforced: the delimiters the sampler counts are set by the chat template regardless, so the cap binds under auto, deepseek and none alike. The default auto already extracts reasoning and is behaviourally identical to deepseek (they differ only in name; the sole branch in the parser is != none). Leave it at the default so the model's own tool-call and channel handling stays in play, and pin deepseek only when a harness needs the thinking kept out of content.

--reasoning-budget on its own forces the closing tag the moment the budget runs out, wherever the model happens to be. When that lands mid-thought the model frequently does not register that it was interrupted: it carries on reasoning, now inside the visible answer. The stop phrase is what prevents that — it gives the model a reason to be finishing.

Two wordings that work

bash
# "qwen" — the string Qwen's own service uses, from their docs
--reasoning-budget-message $'\n\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n'

# "voice" — shorter, in the model's own reasoning voice
--reasoning-budget-message $'\n\nOK, I have enough to answer now.\n'

Wording is model-specific: Qwen note that the ability to act on such a message "is not explicitly trained but emerges naturally", so it is worth trying both on your own workload. Leading and trailing newlines matter — they keep the phrase off whatever half-finished line the cut landed on.

What it measures out to

Measured on the v7-coder IQ4_NL build of this family, served by the same llama.cpp sampler. Three hard questions, temperature 0.6, fixed seed, answer characters with wall time in brackets. Every run answered all three correctly, and thinking length is unchanged by the message in every row:

budgetno message`qwen``voice`
10242284 (35 s)1814 (26 s)1705 (26 s)
204817411 (145 s)1557 (39 s)1673 (39 s)
40961674 (68 s)1538 (67 s)1704 (68 s)

AIME 2024, all 30 problems, budget 4096, -c 32768, vendor sampling:

stop phrasecorrectanswers over 20k charsruns that hit the context wallmean wall
none26/3085159 s
qwen22/301076 s
voice25/301082 s

The phrase halves wall time and all but removes the runaway answers — single rows go from 82,067 characters of answer to 1,655. The accuracy differences are inside the noise at n = 30 (paired: qwen −4 net, voice −1 net, exact binomial p ≈ 0.22 and ≈ 1.0), and the terse "Final Answer:" suffix from the s1 paper (arXiv:2501.19393) is not reproducing the accuracy collapse reported there at this budget.

Per request, instead of per server

The server accepts both as request fields, overriding the command line:

json
{
  "messages": [ ... ],
  "thinking_budget_tokens": 8192,
  "reasoning_budget_message": "\n\nOK, I have enough to answer now.\n"
}

On the raw /completion endpoint the delimiters are not inferred, so they have to be supplied with the budget:

json
{
  "prompt": "...",
  "reasoning_budget_tokens": 8192,
  "reasoning_budget_start_tag": "<|channel>",
  "reasoning_budget_end_tag": "<channel|>",
  "reasoning_budget_message": "\n\nOK, I have enough to answer now.\n"
}

On b10091 the message field must be present on /completion requests even when empty: llama.cpp builds the sequence it forces from message + end_tag inside that field's handler, so omitting it leaves the budget with nothing to force — the sampler logs as though the cap fired while the thinking block stays open.

Rules of thumb

  • —Keep -c several times larger than the budget. A budget equal to the context lets the thinking phase fill the window on its own.
  • —A quarter of the context is a sensible starting point: 8192 at -c 32768.
  • —The budget is per thinking block, not per response — the sampler re-arms when it sees a new opening tag, so a multi-turn agent gets a fresh window each time.

Intended use

A compact (~12–13 GB at Q4_K_M / NVFP4A16, fits a single 12–16 GB GPU) Gemma 4 checkpoint for maximal coding throughput and instruction-following — the code-extreme (x) member of the v7-coder cohort. For the broader LCB-medium lead and HumanEval, use **v7-coder**, which leads those slices (GPQA ≈ 51 on both siblings — neither recovers science).

Inherits Gemma 4's thinking format — serve with the reasoning parser enabled (--reasoning-parser gemma4 on vLLM; --reasoning-format deepseek --reasoning-budget 8192 on llama-server).

Limitations

A research prune, not an official Google release. Expert pruning trades breadth for size: generic_multilingual is de-weighted (0×) and graduate science (GPQA) is a budget axis — at 51.01% it is well below the unpruned 128e (67.17% on the same Q6K run), on par with v7-coder (51.52%); neither sibling recovers science. Quality below ~Q3 / 3-bit degrades on the Gemma 4 MoE — prefer Q4K_M or higher for production. The GGUF and NVFP4A16 formats are provided for deployment but are not separately benchmarked.

Lineage

128e → (v4 → v5 → v6-coder code line) → v7 competence-map rebuild → code4/lcb3 selection + agentic loop-protection force-keep = v7-coderx. Built and evaluated on the omnimergekit toolchain.