datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grug-think
grug-think
grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good.
big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket.
what in box
100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.grug-think-v3-10k
grug-think-v3-10k
v2 brain short. v2 brain useful. but some v2 brain wear office shirt.
"User wants hello world Python. Provide code." short English, yes. grug, no.
v3 tear off office shirt. keep brain meat.
old: User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation.
new: Need Python hello-world. Tiny snippet. No tool. Give code, brief explain.
complex cave different. grug no crush branch into pebble. exact path, error… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v3-10k.grug-3b-train
grug-3b-train
training data for ProCreations/grug-3b.
grug think in grug. grug answer in normal english. never other way round.
what make this one different
old grug model think short always. easy question, short think - good. hard
question, short think - BAD. answer come out worse because grug not do the work.
this set fix that. every fresh example carry difficulty tier, and tier decide
how many word the think get. validator throw away think too short for tier… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-3b-train.grug-27b-train-v21
grug-27b-train-v21
second brain food. this box teach grug-27b v2.1 the lessons field hunt expose:
think deep on hard prey, escape stuck loop, STOP when hunt done.
what in box (2,330 row, ~6.3M token)
slice
rows
teach what
longswe long_task
372
8-12 call debugging arc, failed test cycle, sacred no-tool final summary
longswe stuck_escape
185
same error 3x = approach dead, SWITCH strategy. no more head-bang wall
longswe superlong
120
14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train-v21.grug-35b-v2-train
grug-35b-v2-train
brain food that make grug-35b-v2.
7,106 row, ~14.5M context token, ~2.5M trained token.
what in box
slice
rows
loss
what
SWE agent trajectory (nebius/smith)
2,271
think-only
real agentic coding, grug think, tool call trained, old say-word context only
API tool convo (toolace/glaive/hermes)
935
think-only
multi-turn tool use
fresh code (gpt-5.5)
1,398
full
hard code task, design think, clean answer
fresh math (gpt-5.5)
1,007
full… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train.grug-27b-train
grug-27b-train
brain food that make grug-27b.
7,106 row, ~14.5M context token, ~2.5M trained token.
what in box
slice
rows
loss
what
SWE agent trajectory (nebius/smith)
2,271
think-only
real agentic coding, grug think, tool call trained, old say-word context only
API tool convo (toolace/glaive/hermes)
935
think-only
multi-turn tool use
fresh code (gpt-5.5)
1,398
full
hard code task, design think, clean answer
fresh math (gpt-5.5)
1,007
full
gsm8k… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train.grug-35b-v2-train-v21
grug-35b-v2-train-v21
second brain food. this box teach grug-35b-v2 v2.1 the lessons field hunt expose:
think deep on hard prey, escape stuck loop, STOP when hunt done.
what in box (2,291 row, ~6.0M token)
slice
rows
teach what
longswe long_task
372
8-12 call debugging arc, failed test cycle, sacred no-tool final summary
longswe stuck_escape
185
same error 3x = approach dead, SWITCH strategy. no more head-bang wall
longswe superlong
120
14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train-v21.grug-think-v2-10k
grug-think-v2-10k
grug make second brain box. first box teach short. second box teach dense.
short not whole goal. useful thought per token goal. hard bug need evidence, branch, exact path, risk, verify plan. grug keep all. grug throw grammar padding in fire.
what in box
exactly 10,000 full agent trajectory
62,722 assistant think turn rewritten by DeepSeek-V4-Pro
human say-word unchanged
system/user/tool observation unchanged
tool call and arguments unchanged… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v2-10k.grug-meetingbank-labels
