CoolFace
Modelpublic

ProCreations/grug-v2-9b

sourceHugging Facemitupdated 2mo agoView on Hugging Face
4likes53downloads
Model Card

grug-v2-9b

grug bird keep same club skill. brain voice now actual grug.

July 15, 2026 default-brain audit: 33 fresh neutral prompts, including 15 tool-enabled coding-agent prompts and the three reported planner-English failure shapes, gave reasoning on 33/33, Grug-clean reasoning on 33/33, and zero hidden style instructions. This repo already used original Ornith chat template with no Grug prompt, so grug not replace good weight merely to claim new weight. Tiny depth-only candidate made no meaningful improvement and stayed rejected. Main weights and strong coding/tool scores below remain same; this note records audit.

grug honest release note: old main-branch weights replaced after dialect repair. same repo name, new merged checkpoint. pre-repair rock stays on backup branch pre-dialect-fix-2026-07-13. local cache user should redownload.

old bird sometimes think: "User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation." short, yes. grug, no.

new bird think:

text
Need Python hello-world. Tiny valid snippet enough. Then one-line explain.

dialect fix train only private <think> target. human answer and tool call get no correction loss. joint gate require style improve + every coding/tool score stay whole.

this full merged 9b model. no adapter needed.

whole bird comparison

Same greedy harness, prompts, parser, runtime, and limits. HumanEval 164 tasks; MBPP first 100 sanitized test tasks; card 18 held-out actions; broad 119 held-out actions.

modelHumanEvalMBPPcard validcard strictcard rightbroad validbroad strictbroad right
Ornith 1.0 9B86.676.072.261.161.1100.092.462.2
Grug v1 9B78.776.088.988.988.999.286.691.6
Grug v2 before dialect fix81.177.0100.0100.0100.0100.0100.092.4
Grug v2 corrected82.977.0100.0100.0100.0100.0100.094.1

same-runtime rerun matter. old card number from different vLLM build not mixed into table. exact JSON rock included in results/.

grug family benchmark

Same prompts, parser, runtime, decoding, and limits for both birds. All numbers are percent; bold marks the best result in each column. Ties make both rocks bold.

modelHumanEvalMBPPcard validcard strictcard rightbroad validbroad strictbroad right
Grug v2 9B82.977.0100.0100.0100.0100.0100.094.1
Grug 35B80.588.094.488.994.4100.0100.095.0

dialect-fix capability gate

testbefore dialect fixafterchange
HumanEval pass@1 %81.182.9+1.8
MBPP pass@1 %77.077.0+0.0
card valid tool %100.0100.0+0.0
card strict tool %100.0100.0+0.0
card right tool %100.0100.0+0.0
broad valid tool %100.0100.0+0.0
broad strict tool %100.0100.0+0.0
broad right tool %92.494.1+1.7

valid = parser find offered tool call. strict = exact schema + required args. right tool = expected next action, not merely valid different club.

dialect gate

Separate 90-prompt held-out suite: 50 trivial, 20 moderate, 20 complex. No prompt used for gradient.

measurebeforeafter
dialect-clean trace %1.11100.0
function-word ratio %7.252.44
User asks/wants/... trace890
no tools needed trace700
need to trace40
complex think median word2825

complex median gate protect brain meat: after must keep at least 80% old median and at least 25 word. grug remove grammar, not reasoning branch.

data repair

`grug-think-v3-10k` rewrites only private reasoning from v2 dataset.

  • 10,000 trajectory / 62,722 think turn
  • 11,889 changed trace
  • exact technical anchor retention: 100.0%
  • function-word ratio: 12.2% -> 11.2%
  • matched planner-English patterns after validation: 0
  • every visible answer, tool call, argument, result, system/user message unchanged

correction recipe

  • start public pre-fix ProCreations/grug-v2-9b
  • that checkpoint already carry verifier RL: valid XML, strict argument, right club, closed think reward with rope back to frozen Grug v1
  • rank-8 LoRA, alpha 16, LR 0.0001; final transformer layers 28 onward only
  • adapter delta scale base=0.75; self_attn=-0.470 selected by held-out style screen, then accepted only by whole capability gate
  • 2,440 examples: {"calibration": 640, "coding": 805, "no_tool": 640, "tool": 355}
  • loss scope: private think tokens only; visible answer/tool call masked
  • max context 4,096; 1 epoch; vision tower frozen
  • synthetic calibration covers hello-world, direct multiplication, concise concept explanation
  • checkpoint release only after joint capability + dialect gate

pre-fix weights remain on branch pre-dialect-fix-2026-07-13.

brain + tool shape

Reasoning stays inside <think>...</think>. Native XML tool call stays:

xml
<tool_call>
<function=bash>
<parameter=command>
python -m pytest -q
</parameter>
</function>
</tool_call>

use

python
from transformers import AutoModelForImageTextToText, AutoTokenizer
import torch

name = "ProCreations/grug-v2-9b"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForImageTextToText.from_pretrained(
    name, dtype=torch.bfloat16, device_map="auto")

Need Transformers with Qwen3.5 support. Popular GGUF rocks live at `ProCreations/grug-v2-9b-gguf`.