CoolFace
20 results

grug

marin-community /grug-moe-mix-swarm Grug-MoE Data-Mix Experiments The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below. Fisher-DSP swarm (default) 840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2). Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.tabulartext-generation1K<n<10K1 likes1.8k downloads12d agoHugging FaceProCreations /grug-think grug-think grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good. big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket. what in box 100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.texttext-generation100K<n<1M35 likes378 downloads3mo agoHugging FaceProCreations /grug-27b-v11-title-fix grug-27b-v1.1 title-fix working set Internal dataset for the OpenCode <tool_call> session-title patch. Not a new model version. The trained adapter is merged back into ProCreations/grug-27b-v1.1 in place, then GGUF and MTP are rebuilt into their existing repos. 0 likes269 downloads1mo agoHugging Faceopen-athena /grug-67b-a2b-agentic-sft-training-data Grug 67B agentic SFT training dataset This directory is a local, revision-pinned reconstruction of the exact 29-component mixture consumed by grug_67b_a2b_sft_s3_agentic. The reconstruction has two representations: converted_hf/ contains the readable converted datasets. Each component is checked out at the full Hugging Face commit recorded by the corresponding Marin document artifact. These 29 snapshots contain 77,012 conversations and occupy about 1.67 GB before filesystem… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/grug-67b-a2b-agentic-sft-training-data.0 likes173 downloads19d agoHugging Facelaion /grug-agentic-s3-step1903-repaired-eval-traces Grug 67B repaired-export agentic evaluation traces This dataset contains the final ATIF episode from 587 de-duplicated attempts in a point-in-time snapshot of three active evaluations of laion/grug-67b-a2b-sft-s3-agentic-step1903-repaired. The snapshot was copied on 2026-07-29 at approximately 18:40 UTC. Suite Expected Terminal attempts Scored Mean reward among scored Exported trajectories ID (dev_set_v2) 300 240 209 0.013963 237 SWE-bench Verified 300 126 81 0 126… See the full description on the dataset page: https://huggingface.co/datasets/laion/grug-agentic-s3-step1903-repaired-eval-traces.textn<1K1 likes139 downloads19d agoHugging Facehari31416 /grug-reasoning-data-and-benchmarks Grug Reasoning Datasets and Benchmark Results This repository contains all training datasets, preference pairs, raw model generation logs, and empirical benchmark results for the Grug reasoning research project (spanning both 1.5B and 7B models). Repository Contents ├── data/ │ ├── 1.5b/ │ │ ├── it-1/ # Iteration 1 SFT data and compressed traces │ │ └── it-2/ # Iteration 2 scaled SFT data │ ├── 7b/ │ │ ├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/hari31416/grug-reasoning-data-and-benchmarks.0 likes122 downloads20d agoHugging Face