CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.2k downloads7mo agoHugging Face02marin-community /grug-moe-mix-swarm Grug-MoE Data-Mix Experiments The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below. Fisher-DSP swarm (default) 840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2). Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.tabulartext-generation1K<n<10K1 likes1.8k downloads11d agoHugging Face03grupo-avispa /wildlife_in_irrigation_ponds Wildlife in Irrigation Ponds Dataset Dataset Summary This dataset supports the training and evaluation of object detection models for monitoring irrigation ponds, with the goal of detecting people and animals that have fallen into the water. It comprises synthetically generated images produced using state-of-the-art diffusion models (Z-Image, FLUX), with a real photograph of a target irrigation pond used as the background. The dataset includes four object classes… See the full description on the dataset page: https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds.imageobject-detection1K<n<10K0 likes562 downloads3mo agoHugging Face04dh-unibe /image-text_historisches-grundbuch-basel_xix-xx Dataset Card for image-text_historisches-grundbuch-basel_xix-xx This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 193.409 samples across 1 split(s). The data are transcriptions (automatically generated) from volume 1 of the Historisches Grundbuch of the city of Basel. The entire collection consists of 193.409 pages, of which 135.763 pages are transcribed. The collection with Ground Truth transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_historisches-grundbuch-basel_xix-xx.image100K<n<1M0 likes546 downloads6mo agoHugging Face05ProCreations /grug-think grug-think grug make dataset. dataset make model think like grug. grug think short. short think cheap. cheap think good. big-brain model think 400 token before poke one tool. grug model think 11 word. same tool poke. same work done. many token saved. token = money. grug like money stay in pocket. what in box 100,891 example. every example = full agent conversation: system, user, assistant, tool message. assistant turn always got <think>grug reasoning</think> first… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think.texttext-generation100K<n<1M35 likes378 downloads3mo agoHugging Face06grushaaaaa /indic-multilingual-asr Indic Multilingual ASR Dataset A multilingual ASR dataset covering 13 major Indian languages with 1.1M+ samples. Usage from datasets import load_dataset ds = load_dataset("grushaaaaa/indic-multilingual-asr", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audio1M<n<10M1 likes364 downloads7mo agoHugging Face07cstr /grundwortschatz-voc-de WortUniversum German Vocabulary Database Status: published at cstr/grundwortschatz-voc-de (GPL-3.0). The GPL-3.0 licensing is why the app does not bundle this database: it downloads grundwortschatz-app.db.gz from here on first use. This dataset is the CC-BY-SA→GPL-3.0 re-distribution form of the data built by the WortUniversum pipeline. Rebuild re-push with pipeline/build_hf_datasets.py --upload. Dataset summary A curated, enriched lexical database of 10,450… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-de.tabulartext-classification10K<n<100K0 likes293 downloads1h agoHugging Face08chcaa /grundtvigs-works Grundtvig's Works (Grundtvigs Værker) Grundtvig's Works is a comprehensive digital humanities dataset containing the complete collected writings of Nicolai Frederik Severin Grundtvig (1783-1872) was one of Denmark’s most influential cultural and intellectual figures. As a critical edition, it includes editorial commentary by philologists and is continually updated. The project is scheduled for completion in 2030 and will comprise 1,000 individual works spanning 35,000 pages. The… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/grundtvigs-works.textn<1K4 likes157 downloads1y agoHugging Face09laion /grug-agentic-s3-step1903-repaired-eval-traces Grug 67B repaired-export agentic evaluation traces This dataset contains the final ATIF episode from 587 de-duplicated attempts in a point-in-time snapshot of three active evaluations of laion/grug-67b-a2b-sft-s3-agentic-step1903-repaired. The snapshot was copied on 2026-07-29 at approximately 18:40 UTC. Suite Expected Terminal attempts Scored Mean reward among scored Exported trajectories ID (dev_set_v2) 300 240 209 0.013963 237 SWE-bench Verified 300 126 81 0 126… See the full description on the dataset page: https://huggingface.co/datasets/laion/grug-agentic-s3-step1903-repaired-eval-traces.textn<1K1 likes139 downloads19d agoHugging Face10ProCreations /grug-think-v3-10k grug-think-v3-10k v2 brain short. v2 brain useful. but some v2 brain wear office shirt. "User wants hello world Python. Provide code." short English, yes. grug, no. v3 tear off office shirt. keep brain meat. old: User wants hello world Python. Simple code snippet, no tools needed. Provide code and brief explanation. new: Need Python hello-world. Tiny snippet. No tool. Give code, brief explain. complex cave different. grug no crush branch into pebble. exact path, error… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v3-10k.texttext-generation10K<n<100K12 likes110 downloads2mo agoHugging Face11ProCreations /grug-3b-train grug-3b-train training data for ProCreations/grug-3b. grug think in grug. grug answer in normal english. never other way round. what make this one different old grug model think short always. easy question, short think - good. hard question, short think - BAD. answer come out worse because grug not do the work. this set fix that. every fresh example carry difficulty tier, and tier decide how many word the think get. validator throw away think too short for tier… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-3b-train.texttext-generation1K<n<10K5 likes103 downloads2mo agoHugging Face12marin-community /grug-67b-a2b-snowball-post-sft-evals Grug 67B-A2B Snowball Post-SFT Evaluation Artifacts This repository contains the raw output bundle for the evaluation reported in marin-community/marin issue #7505. It contains 287 score JSON files and 240 sample JSONL files (about 1.4 GB), preserving the layout created by the evaluation jobs. RESULTS.md is the harvested result table. POLICY.md, LAUNCHER.md, launch_baseline.sh, and harvest_s3.py document the evaluation and harvest workflow. The evaluated checkpoint is… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-67b-a2b-snowball-post-sft-evals.text1K<n<10K0 likes93 downloads2mo agoHugging Face13cstr /grundwortschatz-voc-en WortUniversum English Vocabulary Database Status: published at cstr/grundwortschatz-voc-en (CC-BY-SA 4.0). The shipped app asset is assets/grundwortschatz_en.db.gz in the WortUniversum / words-universe repository; this dataset is its CC-BY-SA 4.0 re-distribution form. Rebuild + re-push with pipeline/build_hf_datasets.py --upload. Dataset summary A UK-English lexical database of 11,539 lemmas covering primary-school vocabulary (CEFR-J A1–B2, YLE… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-en.tabulartext-classification10K<n<100K0 likes92 downloads1d agoHugging Face14gruhit-patel /llama-omni-speech-instruct Llama3.2 Omni Speech Instruct Dataset This dataset is created for the sole purpose of enhancing the LLM capability to become multi-modals. This dataset has speech instruction that a model could use to learn and produce the output thus allowing the model to overcome only text input and extends it capabilities towards processing speech command as well. Dataset Details Dataset Description This dataset can be used to train an LLM model to allow adaptibility in… See the full description on the dataset page: https://huggingface.co/datasets/gruhit-patel/llama-omni-speech-instruct.audioquestion-answering10K<n<100K5 likes89 downloads2y agoHugging Face15ProCreations /grug-35b-v2-train grug-35b-v2-train brain food that make grug-35b-v2. 7,106 row, ~14.5M context token, ~2.5M trained token. what in box slice rows loss what SWE agent trajectory (nebius/smith) 2,271 think-only real agentic coding, grug think, tool call trained, old say-word context only API tool convo (toolace/glaive/hermes) 935 think-only multi-turn tool use fresh code (gpt-5.5) 1,398 full hard code task, design think, clean answer fresh math (gpt-5.5) 1,007 full… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train.texttext-generation1K<n<10K3 likes89 downloads2mo agoHugging Face16ProCreations /grug-27b-train-v21 grug-27b-train-v21 second brain food. this box teach grug-27b v2.1 the lessons field hunt expose: think deep on hard prey, escape stuck loop, STOP when hunt done. what in box (2,330 row, ~6.3M token) slice rows teach what longswe long_task 372 8-12 call debugging arc, failed test cycle, sacred no-tool final summary longswe stuck_escape 185 same error 3x = approach dead, SWITCH strategy. no more head-bang wall longswe superlong 120 14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train-v21.texttext-generation1K<n<10K3 likes89 downloads2mo agoHugging Face17grushaaaaa /tts-indian TTS Indian Languages Dataset Speech dataset for Text-to-Speech covering 6 Indian languages, collected and processed from YouTube. Languages & Speakers Speaker Language Gender monihara_bengali Bengali Male munir_kashmiri Kashmiri Male nandini_gujarati Gujarati Female sansri_kannada Kannada Female tamil_pokkisham Tamil Male teluguM Telugu Male Pipeline Audio was collected and processed through these stages: YouTube Download — yt-dlp… See the full description on the dataset page: https://huggingface.co/datasets/grushaaaaa/tts-indian.audio10K<n<100K0 likes82 downloads6mo agoHugging Face18ProCreations /grug-27b-train grug-27b-train brain food that make grug-27b. 7,106 row, ~14.5M context token, ~2.5M trained token. what in box slice rows loss what SWE agent trajectory (nebius/smith) 2,271 think-only real agentic coding, grug think, tool call trained, old say-word context only API tool convo (toolace/glaive/hermes) 935 think-only multi-turn tool use fresh code (gpt-5.5) 1,398 full hard code task, design think, clean answer fresh math (gpt-5.5) 1,007 full gsm8k… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-27b-train.texttext-generation1K<n<10K3 likes76 downloads2mo agoHugging Face19gruhit-patel /libritts_r_train33kaudio10K<n<100K0 likes74 downloads2y agoHugging Face20laion /grug-agentic-s3-step1903-agentic-evals-tracestextn<1K0 likes62 downloads2mo agoHugging Face21ProCreations /grug-35b-v2-train-v21 grug-35b-v2-train-v21 second brain food. this box teach grug-35b-v2 v2.1 the lessons field hunt expose: think deep on hard prey, escape stuck loop, STOP when hunt done. what in box (2,291 row, ~6.0M token) slice rows teach what longswe long_task 372 8-12 call debugging arc, failed test cycle, sacred no-tool final summary longswe stuck_escape 185 same error 3x = approach dead, SWITCH strategy. no more head-bang wall longswe superlong 120 14-20 call… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-35b-v2-train-v21.texttext-generation1K<n<10K2 likes60 downloads2mo agoHugging Face22jalauer /Gruendler2009 Gruendler2009: EEG Obsessive-Compulsive Disorder Estimation Dataset with Flanker Task The dataset Gruendler2009 originates from an EEG experiment that examined the relationship between obsessive–compulsive (OC) symptomatology and error-related brain activity. Participants were 46 undergraduate students selected based on their scores on the Obsessive–Compulsive Inventory-Revised (OCI-R), with groups categorized as high or low OC (a score of 21 being the boundary). The experiment… See the full description on the dataset page: https://huggingface.co/datasets/jalauer/Gruendler2009.textn<1K0 likes59 downloads1y agoHugging Face23DCAgent2 /20260727-222003-grug-agentic-s3-step1903-grug-opencode-id-f947-tracestextn<1K0 likes59 downloads2mo agoHugging Face24Kurmeran /phase2-b2-grubutext100K<n<1M0 likes56 downloads3mo agoHugging Face25dh-unibe /image-text_historisches-grundbuch-basel_xix-xx_train Dataset Card for image-text_historisches-grundbuch-basel_xix-xx_train This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 1962 samples across 1 split(s). The data are transcriptions (Ground Truth) from volume 1 of the Historisches Grundbuch of the city of Basel. The entire collection consists of 193.409 pages, of which 135.763 pages are transcribed. The collection with automatically transcribed data can… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_historisches-grundbuch-basel_xix-xx_train.image1K<n<10K0 likes55 downloads6mo agoHugging Face26GRuri /ssgentext1K<n<10K0 likes53 downloads11mo agoHugging Face27umsa-v1 /dataset_parafraseado_grupo1 🎓 Contexto Académico Maestría en Inteligencia Artificial y Data Science para la Transformación de Negocios 🏛️ Institución: Postgrado de Informática 📚 Módulo: Modelamiento de Datos II 👨‍🏫 Docente: Prof. Anvi Alex Eponon 📅 Año: 2024 📚 dataset_parafraseado_grupo1 Dataset académico para entrenamiento de modelos especializados en RGPD/GDPR Grupo 1 | Modelamiento de Datos II 📖 Descripción dataset_parafraseado_grupo1 es un conjunto de datos… See the full description on the dataset page: https://huggingface.co/datasets/umsa-v1/dataset_parafraseado_grupo1.texttext-generationn<1K0 likes49 downloads8mo agoHugging Face28Kurmeran /phase2-b-grubutext100K<n<1M1 likes48 downloads3mo agoHugging Face29ProCreations /grug-think-v2-10k grug-think-v2-10k grug make second brain box. first box teach short. second box teach dense. short not whole goal. useful thought per token goal. hard bug need evidence, branch, exact path, risk, verify plan. grug keep all. grug throw grammar padding in fire. what in box exactly 10,000 full agent trajectory 62,722 assistant think turn rewritten by DeepSeek-V4-Pro human say-word unchanged system/user/tool observation unchanged tool call and arguments unchanged… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/grug-think-v2-10k.texttext-generation10K<n<100K4 likes40 downloads2mo agoHugging Face30grupo4-bisite /mbpp_finaltext1K<n<10K0 likes36 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.