gru
Datasets
All datasets matching “gru”indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
grug-moe-mix-swarm
Grug-MoE Data-Mix Experiments
The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below.
Fisher-DSP swarm (default)
840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2).
Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress
mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.BIRDeep_AudioAnnotations
BIRDeep Audio Annotations
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification.
The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/GrunCrow/BIRDeep_AudioAnnotations.wildlife_in_irrigation_ponds
Wildlife in Irrigation Ponds Dataset
Dataset Summary
This dataset supports the training and evaluation of object detection models for monitoring irrigation ponds, with the goal of detecting people and animals that have fallen into the water. It comprises synthetically generated images produced using state-of-the-art diffusion models (Z-Image, FLUX), with a real photograph of a target irrigation pond used as the background.
The dataset includes four object classes… See the full description on the dataset page: https://huggingface.co/datasets/grupo-avispa/wildlife_in_irrigation_ponds.image-text_historisches-grundbuch-basel_xix-xx
Dataset Card for image-text_historisches-grundbuch-basel_xix-xx
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 193.409 samples across 1 split(s). The data are transcriptions (automatically generated) from volume 1 of the Historisches Grundbuch of the city of Basel. The entire collection consists of 193.409 pages, of which 135.763 pages are transcribed. The collection with Ground Truth transcriptions… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_historisches-grundbuch-basel_xix-xx.GenManip-Assets-GRUtopiaTableSubset
