datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
grug-moe-mix-swarm
Grug-MoE Data-Mix Experiments
The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below.
Fisher-DSP swarm (default)
840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2).
Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress
mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.BIRDeep_AudioAnnotations
BIRDeep Audio Annotations
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification.
The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/GrunCrow/BIRDeep_AudioAnnotations.wildlife_in_irrigation_ponds
Wildlife in Irrigation Ponds Dataset
Dataset Summary
This dataset supports the training and evaluation of object detection models for monitoring irrigation ponds, with the goal of detecting people and animals that have fallen into the water. It comprises synthetically generated images produced using state-of-the-art diffusion models (Z-Image, FLUX), with a real photograph of a target irrigation pond used as the background.
The dataset includes four object classes… See the full description on the dataset page: https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds.image-text_historisches-grundbuch-basel_xix-xx
Dataset Card for image-text_historisches-grundbuch-basel_xix-xx
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 193.409 samples across 1 split(s). The data are transcriptions (automatically generated) from volume 1 of the Historisches Grundbuch of the city of Basel. The entire collection consists of 193.409 pages, of which 135.763 pages are transcribed. The collection with Ground Truth transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_historisches-grundbuch-basel_xix-xx.GenManip-Assets-GRUtopiaTableSubsetgrug-think
grug-think
grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good.
big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket.
what in box
100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.indic-multilingual-asr
Indic Multilingual ASR Dataset
A multilingual ASR dataset covering 13 major Indian languages with 1.1M+ samples.
Usage
from datasets import load_dataset
ds = load_dataset("grushaaaaa/indic-multilingual-asr", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
grundwortschatz-voc-de
WortUniversum German Vocabulary Database
Status: published at
cstr/grundwortschatz-voc-de
(GPL-3.0). The GPL-3.0 licensing is why the app does not bundle this database:
it downloads grundwortschatz-app.db.gz from here on first use. This dataset
is the CC-BY-SA→GPL-3.0 re-distribution form of the data built by the
WortUniversum pipeline. Rebuild
re-push with pipeline/build_hf_datasets.py --upload.
Dataset summary
A curated, enriched lexical database of 10,450… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-de.grug-27b-v11-title-fix
grug-27b-v1.1 title-fix working set
Internal dataset for the OpenCode <tool_call> session-title patch.
Not a new model version. The trained adapter is merged back into
ProCreations/grug-27b-v1.1 in place, then GGUF and MTP are rebuilt
into their existing repos.
grug-67b-a2b-agentic-sft-training-data
Grug 67B agentic SFT training dataset
This directory is a local, revision-pinned reconstruction of the exact 29-component mixture consumed by grug_67b_a2b_sft_s3_agentic.
The reconstruction has two representations:
converted_hf/ contains the readable converted datasets. Each component is checked out at the full Hugging Face commit recorded by the corresponding Marin document artifact. These 29 snapshots contain 77,012 conversations and occupy about 1.67 GB before filesystem… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/grug-67b-a2b-agentic-sft-training-data.grundtvigs-works
Grundtvig's Works (Grundtvigs Værker)
Grundtvig's Works is a comprehensive digital humanities dataset containing the complete collected writings of
Nicolai Frederik Severin Grundtvig (1783-1872) was one of Denmark’s most influential cultural and intellectual figures.
As a critical edition, it includes editorial commentary by philologists and is continually updated.
The project is scheduled for completion in 2030 and will comprise 1,000 individual works spanning 35,000 pages. The… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/grundtvigs-works.grug-agentic-s3-step1903-repaired-eval-traces
Grug 67B repaired-export agentic evaluation traces
This dataset contains the final ATIF episode from 587 de-duplicated attempts in
a point-in-time snapshot of three active evaluations of
laion/grug-67b-a2b-sft-s3-agentic-step1903-repaired.
The snapshot was copied on 2026-07-29 at approximately 18:40 UTC.
Suite
Expected
Terminal attempts
Scored
Mean reward among scored
Exported trajectories
ID (dev_set_v2)
300
240
209
0.013963
237
SWE-bench Verified
300
126
81
0
126… See the full description on the dataset page: https://huggingface.co/datasets/laion/grug-agentic-s3-step1903-repaired-eval-traces.grug-reasoning-data-and-benchmarks
Grug Reasoning Datasets and Benchmark Results
This repository contains all training datasets, preference pairs, raw model generation logs, and empirical benchmark results for the Grug reasoning research project (spanning both 1.5B and 7B models).
Repository Contents
├── data/
│ ├── 1.5b/
│ │ ├── it-1/ # Iteration 1 SFT data and compressed traces
│ │ └── it-2/ # Iteration 2 scaled SFT data
│ ├── 7b/
│ │ ├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/hari31416/grug-reasoning-data-and-benchmarks.grug-3b-train
grug-3b-train
training data for ProCreations/grug-3b.
grug think in grug. grug answer in normal english. never other way round.
what make this one different
old grug model think short always. easy question, short think - good. hard
question, short think - BAD. answer come out worse because grug not do the work.
this set fix that. every fresh example carry difficulty tier, and tier decide
how many word the think get. validator throw away think too short for tier… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-3b-train.grug-think-v3-10k
grug-think-v3-10k
v2 brain short. v2 brain useful. but some v2 brain wear office shirt.
"User wants hello world Python. Provide code." short English, yes. grug, no.
v3 tear off office shirt. keep brain meat.
old: User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation.
new: Need Python hello-world. Tiny snippet. No tool. Give code, brief explain.
complex cave different. grug no crush branch into pebble. exact path, error… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v3-10k.grug-67b-a2b-snowball-post-sft-evals
Grug 67B-A2B Snowball Post-SFT Evaluation Artifacts
This repository contains the raw output bundle for the evaluation reported in
marin-community/marin issue #7505.
It contains 287 score JSON files and 240 sample JSONL files (about 1.4 GB),
preserving the layout created by the evaluation jobs.
RESULTS.md is the harvested result table. POLICY.md, LAUNCHER.md,
launch_baseline.sh, and harvest_s3.py document the evaluation and harvest
workflow. The evaluated checkpoint is… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-67b-a2b-snowball-post-sft-evals.llama-omni-speech-instruct
Llama3.2 Omni Speech Instruct Dataset
This dataset is created for the sole purpose of enhancing the LLM capability to become multi-modals. This dataset has speech instruction
that a model could use to learn and produce the output thus allowing the model to overcome only text input and extends it capabilities
towards processing speech command as well.
Dataset Details
Dataset Description
This dataset can be used to train an LLM model to allow adaptibility in… See the full description on the dataset page: https://huggingface.co/datasets/gruhit-patel/llama-omni-speech-instruct.grug-35b-v2-train
grug-35b-v2-train
brain food that make grug-35b-v2.
7,106 row, ~14.5M context token, ~2.5M trained token.
what in box
slice
rows
loss
what
SWE agent trajectory (nebius/smith)
2,271
think-only
real agentic coding, grug think, tool call trained, old say-word context only
API tool convo (toolace/glaive/hermes)
935
think-only
multi-turn tool use
fresh code (gpt-5.5)
1,398
full
hard code task, design think, clean answer
fresh math (gpt-5.5)
1,007
full… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train.grug-27b-train-v21
grug-27b-train-v21
second brain food. this box teach grug-27b v2.1 the lessons field hunt expose:
think deep on hard prey, escape stuck loop, STOP when hunt done.
what in box (2,330 row, ~6.3M token)
slice
rows
teach what
longswe long_task
372
8-12 call debugging arc, failed test cycle, sacred no-tool final summary
longswe stuck_escape
185
same error 3x = approach dead, SWITCH strategy. no more head-bang wall
longswe superlong
120
14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train-v21.tts-indian
TTS Indian Languages Dataset
Speech dataset for Text-to-Speech covering 6 Indian languages, collected and processed from YouTube.
Languages & Speakers
Speaker
Language
Gender
monihara_bengali
Bengali
Male
munir_kashmiri
Kashmiri
Male
nandini_gujarati
Gujarati
Female
sansri_kannada
Kannada
Female
tamil_pokkisham
Tamil
Male
teluguM
Telugu
Male
Pipeline
Audio was collected and processed through these stages:
YouTube Download — yt-dlp… See the full description on the dataset page: https://huggingface.co/datasets/grushaaaaa/tts-indian.grundwortschatz-voc-en
WortUniversum English Vocabulary Database
Status: published at
cstr/grundwortschatz-voc-en
(CC-BY-SA 4.0). The shipped app asset is assets/grundwortschatz_en.db.gz in
the WortUniversum / words-universe
repository; this dataset is its CC-BY-SA 4.0 re-distribution form. Rebuild +
re-push with pipeline/build_hf_datasets.py --upload.
Dataset summary
A UK-English lexical database of 11,539 lemmas covering primary-school
vocabulary (CEFR-J A1–B2, YLE… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-en.grug-27b-train
grug-27b-train
brain food that make grug-27b.
7,106 row, ~14.5M context token, ~2.5M trained token.
what in box
slice
rows
loss
what
SWE agent trajectory (nebius/smith)
2,271
think-only
real agentic coding, grug think, tool call trained, old say-word context only
API tool convo (toolace/glaive/hermes)
935
think-only
multi-turn tool use
fresh code (gpt-5.5)
1,398
full
hard code task, design think, clean answer
fresh math (gpt-5.5)
1,007
full
gsm8k… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train.libritts_r_train33krot_gruen_sort_20260724_094754This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sara2369/rot_gruen_sort_20260724_094754.few-shot-grumpy-cat
Citation
@article{DBLP:journals/corr/abs-2101-04775,
author = {Bingchen Liu and
Yizhe Zhu and
Kunpeng Song and
Ahmed Elgammal},
title = {Towards Faster and Stabilized {GAN} Training for High-fidelity Few-shot
Image Synthesis},
journal = {CoRR},
volume = {abs/2101.04775},
year = {2021},
url = {https://arxiv.org/abs/2101.04775},
eprinttype = {arXiv},
eprint = {2101.04775},
timestamp… See the full description on the dataset page: https://huggingface.co/datasets/huggan/few-shot-grumpy-cat.Gruendler2009
Gruendler2009: EEG Obsessive-Compulsive Disorder Estimation Dataset with Flanker Task
The dataset Gruendler2009 originates from an EEG experiment that examined the relationship between obsessive–compulsive (OC) symptomatology and error-related brain activity. Participants were 46 undergraduate students selected based on their scores on the Obsessive–Compulsive Inventory-Revised (OCI-R), with groups categorized as high or low OC (a score of 21 being the boundary).
The experiment… See the full description on the dataset page: https://huggingface.co/datasets/jalauer/Gruendler2009.grug-agentic-s3-step1903-agentic-evals-tracesgrug-35b-v2-train-v21
grug-35b-v2-train-v21
second brain food. this box teach grug-35b-v2 v2.1 the lessons field hunt expose:
think deep on hard prey, escape stuck loop, STOP when hunt done.
what in box (2,291 row, ~6.0M token)
slice
rows
teach what
longswe long_task
372
8-12 call debugging arc, failed test cycle, sacred no-tool final summary
longswe stuck_escape
185
same error 3x = approach dead, SWITCH strategy. no more head-bang wall
longswe superlong
120
14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train-v21.grustnogram
Dataset Card for Grustnogram
Dataset Summary
This dataset contains 597,704 posts from Grustnogram.ru, a Russian "emotional network" similar to Instagram but with a distinctive black and white filter aesthetic and dark atmosphere. The dataset includes 542,917 image posts with associated metadata and 54,787 anonymous text-only posts.
Languages
The dataset is primarily in Russian (ru).
Dataset Structure
Data Splits
The dataset is divided into… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/grustnogram.
