CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.2k downloads7mo agoHugging Face02marin-community /grug-moe-mix-swarm Grug-MoE Data-Mix Experiments The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below. Fisher-DSP swarm (default) 840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2). Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.tabulartext-generation1K<n<10K1 likes1.9k downloads10d agoHugging Face03GrunCrow /BIRDeep_AudioAnnotations BIRDeep Audio Annotations The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification. The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/GrunCrow/BIRDeep_AudioAnnotations.audioaudio-classificationn<1K2 likes884 downloads11mo agoHugging Face04grupo-avispa /wildlife_in_irrigation_ponds Wildlife in Irrigation Ponds Dataset Dataset Summary This dataset supports the training and evaluation of object detection models for monitoring irrigation ponds, with the goal of detecting people and animals that have fallen into the water. It comprises synthetically generated images produced using state-of-the-art diffusion models (Z-Image, FLUX), with a real photograph of a target irrigation pond used as the background. The dataset includes four object classes… See the full description on the dataset page: https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds.imageobject-detection1K<n<10K0 likes560 downloads3mo agoHugging Face05dh-unibe /image-text_historisches-grundbuch-basel_xix-xx Dataset Card for image-text_historisches-grundbuch-basel_xix-xx This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 193.409 samples across 1 split(s). The data are transcriptions (automatically generated) from volume 1 of the Historisches Grundbuch of the city of Basel. The entire collection consists of 193.409 pages, of which 135.763 pages are transcribed. The collection with Ground Truth transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_historisches-grundbuch-basel_xix-xx.image100K<n<1M0 likes559 downloads6mo agoHugging Face06Axi404 /GenManip-Assets-GRUtopiaTableSubset0 likes540 downloads8mo agoHugging Face07ProCreations /grug-think grug-think grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good. big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket. what in box 100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.texttext-generation100K<n<1M35 likes370 downloads3mo agoHugging Face08grushaaaaa /indic-multilingual-asr Indic Multilingual ASR Dataset A multilingual ASR dataset covering 13 major Indian languages with 1.1M+ samples. Usage from datasets import load_dataset ds = load_dataset("grushaaaaa/indic-multilingual-asr", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audio1M<n<10M1 likes359 downloads7mo agoHugging Face09cstr /grundwortschatz-voc-de WortUniversum German Vocabulary Database Status: published at cstr/grundwortschatz-voc-de (GPL-3.0). The GPL-3.0 licensing is why the app does not bundle this database: it downloads grundwortschatz-app.db.gz from here on first use. This dataset is the CC-BY-SA→GPL-3.0 re-distribution form of the data built by the WortUniversum pipeline. Rebuild re-push with pipeline/build_hf_datasets.py --upload. Dataset summary A curated, enriched lexical database of 10,450… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-de.tabulartext-classification10K<n<100K0 likes275 downloads7h agoHugging Face10ProCreations /grug-27b-v11-title-fix grug-27b-v1.1 title-fix working set Internal dataset for the OpenCode <tool_call> session-title patch. Not a new model version. The trained adapter is merged back into ProCreations/grug-27b-v1.1 in place, then GGUF and MTP are rebuilt into their existing repos. 0 likes265 downloads1mo agoHugging Face11open-athena /grug-67b-a2b-agentic-sft-training-data Grug 67B agentic SFT training dataset This directory is a local, revision-pinned reconstruction of the exact 29-component mixture consumed by grug_67b_a2b_sft_s3_agentic. The reconstruction has two representations: converted_hf/ contains the readable converted datasets. Each component is checked out at the full Hugging Face commit recorded by the corresponding Marin document artifact. These 29 snapshots contain 77,012 conversations and occupy about 1.67 GB before filesystem… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/grug-67b-a2b-agentic-sft-training-data.0 likes169 downloads18d agoHugging Face12chcaa /grundtvigs-works Grundtvig's Works (Grundtvigs Værker) Grundtvig's Works is a comprehensive digital humanities dataset containing the complete collected writings of Nicolai Frederik Severin Grundtvig (1783-1872) was one of Denmark’s most influential cultural and intellectual figures. As a critical edition, it includes editorial commentary by philologists and is continually updated. The project is scheduled for completion in 2030 and will comprise 1,000 individual works spanning 35,000 pages. The… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/grundtvigs-works.textn<1K4 likes161 downloads1y agoHugging Face13laion /grug-agentic-s3-step1903-repaired-eval-traces Grug 67B repaired-export agentic evaluation traces This dataset contains the final ATIF episode from 587 de-duplicated attempts in a point-in-time snapshot of three active evaluations of laion/grug-67b-a2b-sft-s3-agentic-step1903-repaired. The snapshot was copied on 2026-07-29 at approximately 18:40 UTC. Suite Expected Terminal attempts Scored Mean reward among scored Exported trajectories ID (dev_set_v2) 300 240 209 0.013963 237 SWE-bench Verified 300 126 81 0 126… See the full description on the dataset page: https://huggingface.co/datasets/laion/grug-agentic-s3-step1903-repaired-eval-traces.textn<1K1 likes139 downloads18d agoHugging Face14hari31416 /grug-reasoning-data-and-benchmarks Grug Reasoning Datasets and Benchmark Results This repository contains all training datasets, preference pairs, raw model generation logs, and empirical benchmark results for the Grug reasoning research project (spanning both 1.5B and 7B models). Repository Contents ├── data/ │ ├── 1.5b/ │ │ ├── it-1/ # Iteration 1 SFT data and compressed traces │ │ └── it-2/ # Iteration 2 scaled SFT data │ ├── 7b/ │ │ ├── sft/… See the full description on the dataset page: https://huggingface.co/datasets/hari31416/grug-reasoning-data-and-benchmarks.0 likes122 downloads19d agoHugging Face15ProCreations /grug-3b-train grug-3b-train training data for ProCreations/grug-3b. grug think in grug. grug answer in normal english. never other way round. what make this one different old grug model think short always. easy question, short think - good. hard question, short think - BAD. answer come out worse because grug not do the work. this set fix that. every fresh example carry difficulty tier, and tier decide how many word the think get. validator throw away think too short for tier… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-3b-train.texttext-generation1K<n<10K5 likes110 downloads2mo agoHugging Face16ProCreations /grug-think-v3-10k grug-think-v3-10k v2 brain short. v2 brain useful. but some v2 brain wear office shirt. "User wants hello world Python. Provide code." short English, yes. grug, no. v3 tear off office shirt. keep brain meat. old: User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation. new: Need Python hello-world. Tiny snippet. No tool. Give code, brief explain. complex cave different. grug no crush branch into pebble. exact path, error… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v3-10k.texttext-generation10K<n<100K12 likes109 downloads2mo agoHugging Face17marin-community /grug-67b-a2b-snowball-post-sft-evals Grug 67B-A2B Snowball Post-SFT Evaluation Artifacts This repository contains the raw output bundle for the evaluation reported in marin-community/marin issue #7505. It contains 287 score JSON files and 240 sample JSONL files (about 1.4 GB), preserving the layout created by the evaluation jobs. RESULTS.md is the harvested result table. POLICY.md, LAUNCHER.md, launch_baseline.sh, and harvest_s3.py document the evaluation and harvest workflow. The evaluated checkpoint is… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-67b-a2b-snowball-post-sft-evals.text1K<n<10K0 likes90 downloads2mo agoHugging Face18gruhit-patel /llama-omni-speech-instruct Llama3.2 Omni Speech Instruct Dataset This dataset is created for the sole purpose of enhancing the LLM capability to become multi-modals. This dataset has speech instruction that a model could use to learn and produce the output thus allowing the model to overcome only text input and extends it capabilities towards processing speech command as well. Dataset Details Dataset Description This dataset can be used to train an LLM model to allow adaptibility in… See the full description on the dataset page: https://huggingface.co/datasets/gruhit-patel/llama-omni-speech-instruct.audioquestion-answering10K<n<100K5 likes89 downloads2y agoHugging Face19ProCreations /grug-35b-v2-train grug-35b-v2-train brain food that make grug-35b-v2. 7,106 row, ~14.5M context token, ~2.5M trained token. what in box slice rows loss what SWE agent trajectory (nebius/smith) 2,271 think-only real agentic coding, grug think, tool call trained, old say-word context only API tool convo (toolace/glaive/hermes) 935 think-only multi-turn tool use fresh code (gpt-5.5) 1,398 full hard code task, design think, clean answer fresh math (gpt-5.5) 1,007 full… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train.texttext-generation1K<n<10K3 likes87 downloads2mo agoHugging Face20ProCreations /grug-27b-train-v21 grug-27b-train-v21 second brain food. this box teach grug-27b v2.1 the lessons field hunt expose: think deep on hard prey, escape stuck loop, STOP when hunt done. what in box (2,330 row, ~6.3M token) slice rows teach what longswe long_task 372 8-12 call debugging arc, failed test cycle, sacred no-tool final summary longswe stuck_escape 185 same error 3x = approach dead, SWITCH strategy. no more head-bang wall longswe superlong 120 14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train-v21.texttext-generation1K<n<10K3 likes87 downloads2mo agoHugging Face21grushaaaaa /tts-indian TTS Indian Languages Dataset Speech dataset for Text-to-Speech covering 6 Indian languages, collected and processed from YouTube. Languages & Speakers Speaker Language Gender monihara_bengali Bengali Male munir_kashmiri Kashmiri Male nandini_gujarati Gujarati Female sansri_kannada Kannada Female tamil_pokkisham Tamil Male teluguM Telugu Male Pipeline Audio was collected and processed through these stages: YouTube Download — yt-dlp… See the full description on the dataset page: https://huggingface.co/datasets/grushaaaaa/tts-indian.audio10K<n<100K0 likes80 downloads6mo agoHugging Face22cstr /grundwortschatz-voc-en WortUniversum English Vocabulary Database Status: published at cstr/grundwortschatz-voc-en (CC-BY-SA 4.0). The shipped app asset is assets/grundwortschatz_en.db.gz in the WortUniversum / words-universe repository; this dataset is its CC-BY-SA 4.0 re-distribution form. Rebuild + re-push with pipeline/build_hf_datasets.py --upload. Dataset summary A UK-English lexical database of 11,539 lemmas covering primary-school vocabulary (CEFR-J A1–B2, YLE… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-en.tabulartext-classification10K<n<100K0 likes75 downloads7h agoHugging Face23ProCreations /grug-27b-train grug-27b-train brain food that make grug-27b. 7,106 row, ~14.5M context token, ~2.5M trained token. what in box slice rows loss what SWE agent trajectory (nebius/smith) 2,271 think-only real agentic coding, grug think, tool call trained, old say-word context only API tool convo (toolace/glaive/hermes) 935 think-only multi-turn tool use fresh code (gpt-5.5) 1,398 full hard code task, design think, clean answer fresh math (gpt-5.5) 1,007 full gsm8k… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train.texttext-generation1K<n<10K3 likes74 downloads2mo agoHugging Face24gruhit-patel /libritts_r_train33kaudio10K<n<100K0 likes73 downloads2y agoHugging Face25sara2369 /rot_gruen_sort_20260724_094754This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/sara2369/rot_gruen_sort_20260724_094754.tabularrobotics10K<n<100K0 likes72 downloads2mo agoHugging Face26huggan /few-shot-grumpy-cat Citation @article{DBLP:journals/corr/abs-2101-04775, author = {Bingchen Liu and Yizhe Zhu and Kunpeng Song and Ahmed Elgammal}, title = {Towards Faster and Stabilized {GAN} Training for High-fidelity Few-shot Image Synthesis}, journal = {CoRR}, volume = {abs/2101.04775}, year = {2021}, url = {https://arxiv.org/abs/2101.04775}, eprinttype = {arXiv}, eprint = {2101.04775}, timestamp… See the full description on the dataset page: https://huggingface.co/datasets/huggan/few-shot-grumpy-cat.imagen<1K0 likes65 downloads4y agoHugging Face27jalauer /Gruendler2009 Gruendler2009: EEG Obsessive-Compulsive Disorder Estimation Dataset with Flanker Task The dataset Gruendler2009 originates from an EEG experiment that examined the relationship between obsessive–compulsive (OC) symptomatology and error-related brain activity. Participants were 46 undergraduate students selected based on their scores on the Obsessive–Compulsive Inventory-Revised (OCI-R), with groups categorized as high or low OC (a score of 21 being the boundary). The experiment… See the full description on the dataset page: https://huggingface.co/datasets/jalauer/Gruendler2009.textn<1K0 likes63 downloads1y agoHugging Face28laion /grug-agentic-s3-step1903-agentic-evals-tracestextn<1K0 likes62 downloads2mo agoHugging Face29ProCreations /grug-35b-v2-train-v21 grug-35b-v2-train-v21 second brain food. this box teach grug-35b-v2 v2.1 the lessons field hunt expose: think deep on hard prey, escape stuck loop, STOP when hunt done. what in box (2,291 row, ~6.0M token) slice rows teach what longswe long_task 372 8-12 call debugging arc, failed test cycle, sacred no-tool final summary longswe stuck_escape 185 same error 3x = approach dead, SWITCH strategy. no more head-bang wall longswe superlong 120 14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train-v21.texttext-generation1K<n<10K2 likes61 downloads2mo agoHugging Face30nyuuzyou /grustnogram Dataset Card for Grustnogram Dataset Summary This dataset contains 597,704 posts from Grustnogram.ru, a Russian "emotional network" similar to Instagram but with a distinctive black and white filter aesthetic and dark atmosphere. The dataset includes 542,917 image posts with associated metadata and 54,787 anonymous text-only posts. Languages The dataset is primarily in Russian (ru). Dataset Structure Data Splits The dataset is divided into… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/grustnogram.imageimage-classification100K<n<1M1 likes59 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.