ProCreations/grug-v2-9b
grug-v2-9b
grug bird keep same club skill. brain voice now actual grug.
July 15, 2026 default-brain audit: 33 fresh neutral prompts, including 15 tool-enabled coding-agent prompts and the three reported planner-English failure shapes, gave reasoning on 33/33, Grug-clean reasoning on 33/33, and zero hidden style instructions. This repo already used original Ornith chat template with no Grug prompt, so grug not replace good weight merely to claim new weight. Tiny depth-only candidate made no meaningful improvement and stayed rejected. Main weights and strong coding/tool scores below remain same; this note records audit.
grug honest release note: old main-branch weights replaced after dialect repair. same repo name, new merged checkpoint. pre-repair rock stays on backup branch pre-dialect-fix-2026-07-13. local cache user should redownload.
old bird sometimes think: "User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation." short, yes. grug, no.
new bird think:
Need Python hello-world. Tiny valid snippet enough. Then one-line explain.dialect fix train only private <think> target. human answer and tool call get no correction loss. joint gate require style improve + every coding/tool score stay whole.
this full merged 9b model. no adapter needed.
whole bird comparison
Same greedy harness, prompts, parser, runtime, and limits. HumanEval 164 tasks; MBPP first 100 sanitized test tasks; card 18 held-out actions; broad 119 held-out actions.
same-runtime rerun matter. old card number from different vLLM build not mixed into table. exact JSON rock included in results/.
grug family benchmark
Same prompts, parser, runtime, decoding, and limits for both birds. All numbers are percent; bold marks the best result in each column. Ties make both rocks bold.
dialect-fix capability gate
valid = parser find offered tool call. strict = exact schema + required args. right tool = expected next action, not merely valid different club.
dialect gate
Separate 90-prompt held-out suite: 50 trivial, 20 moderate, 20 complex. No prompt used for gradient.
complex median gate protect brain meat: after must keep at least 80% old median and at least 25 word. grug remove grammar, not reasoning branch.
data repair
`grug-think-v3-10k` rewrites only private reasoning from v2 dataset.
- 10,000 trajectory / 62,722 think turn
- 11,889 changed trace
- exact technical anchor retention: 100.0%
- function-word ratio: 12.2% -> 11.2%
- matched planner-English patterns after validation: 0
- every visible answer, tool call, argument, result, system/user message unchanged
correction recipe
- start public pre-fix
ProCreations/grug-v2-9b - that checkpoint already carry verifier RL: valid XML, strict argument, right club, closed think reward with rope back to frozen Grug v1
- rank-8 LoRA, alpha 16, LR 0.0001; final transformer layers 28 onward only
- adapter delta scale base=0.75; self_attn=-0.470 selected by held-out style screen, then accepted only by whole capability gate
- 2,440 examples: {"calibration": 640, "coding": 805, "no_tool": 640, "tool": 355}
- loss scope: private think tokens only; visible answer/tool call masked
- max context 4,096; 1 epoch; vision tower frozen
- synthetic calibration covers hello-world, direct multiplication, concise concept explanation
- checkpoint release only after joint capability + dialect gate
pre-fix weights remain on branch pre-dialect-fix-2026-07-13.
brain + tool shape
Reasoning stays inside <think>...</think>. Native XML tool call stays:
<tool_call>
<function=bash>
<parameter=command>
python -m pytest -q
</parameter>
</function>
</tool_call>use
from transformers import AutoModelForImageTextToText, AutoTokenizer
import torch
name = "ProCreations/grug-v2-9b"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForImageTextToText.from_pretrained(
name, dtype=torch.bfloat16, device_map="auto")Need Transformers with Qwen3.5 support. Popular GGUF rocks live at `ProCreations/grug-v2-9b-gguf`.
