CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GEM /wiki_auto_asset_turk Dataset Card for GEM/wiki_auto_asset_turk Link to Main Data Card You can find the main data card on the GEM Website. Dataset Summary WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.text100K<n<1M8 likes9.4k downloads2y agoHugging Face02Mechanistic-Anomaly-Detection /gemma2-jailbreakstext10K<n<100K2 likes5.6k downloads2y agoHugging Face03ioi-leaderboard /ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limittextn<1K0 likes2.2k downloads2y agoHugging Face04ghanaopenai /twi-health-asr-gemini-500hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset Gemini (500 hours) A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-health-asr-gemini-500hrs.audioautomatic-speech-recognition10K<n<100K0 likes2k downloads2mo agoHugging Face05Nilaksh404 /gemini-2_5-protextn<1K0 likes1.7k downloads10mo agoHugging Face06shb777 /gemini-flash-2.0-speech 🎙️ Gemini Flash 2.0 Speech Dataset This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English. 🏅 #1 Trending Audio Dataset in Feb 2025 🏅 Used in training of Kokoro TTS and LLaSA 1B 〽️ Stats Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours) Average duration: 10.83 seconds Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.audiotext-to-speech10K<n<100K60 likes1.3k downloads1y agoHugging Face07credi-net /CDB_DEC2024-CochranSampled_Gemma-300m_Embtext100M<n<1B1 likes1.1k downloads2mo agoHugging Face08prism-vlm /gemini_public_mmr1 PRISM Public SFT Data Overview PRISM Public SFT Data is the public supervised fine-tuning data collection used in the PRISM project.PRISM studies the distributional drift problem in the standard SFT → RLVR post-training pipeline for large multimodal models. Before the distribution alignment and RLVR stages, we first use large-scale public multimodal demonstrations to obtain a broad SFT initialization. This dataset serves as the public SFT data source for the… See the full description on the dataset page: https://huggingface.co/datasets/prism-vlm/gemini_public_mmr1.textimage-to-text1M<n<10M2 likes941 downloads5mo agoHugging Face09Nilaksh404 /gemini-3-pro-previewdocumentn<1K0 likes887 downloads10mo agoHugging Face10lightonai /ms-marco-en-bge-gemma ms-marco-en-bge This dataset contains the MS MARCO dataset with negatives mined using ColBERT and then scored by bge-reranker-v2-gemma. It can be used to train a retrieval model using knowledge distillation, for example using PyLate. knowledge distillation To fine-tune a model using knowledge distillation loss we will need three distinct file: Datasetsfrom datasets import load_dataset train = load_dataset( "lightonai/ms-marco-en-gemma", "train"… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/ms-marco-en-bge-gemma.textfeature-extraction10M<n<100M13 likes886 downloads1y agoHugging Face11tomg-group-umd /gemstones_data_order_parallelGemstones Training Dataset - Parallel workers sharded version This data is a reprocessed version of the first 1B rows of the Dolma v1.7 dataset (https://huggingface.co/datasets/allenai/dolma). The data is encoded using the Pythia tokenizer: https://huggingface.co/EleutherAI/pythia-160m Disclaimer: this is an approximation of the dataset used to train the Gemstones model suite. Due to the randomized and sharded nature of the distributed training code, the only way to perfectly reproduce the… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/gemstones_data_order_parallel.100M<n<1B0 likes879 downloads1y agoHugging Face12TAUR-Lab /Taur_CoT_Analysis_Project___google__gemini-1.5-flash-001text10K<n<100K0 likes848 downloads2y agoHugging Face13juiceb0xc0de /gemma-4-e2b-atlas image1M<n<10M4 likes810 downloads10d agoHugging Face14kalomaze /glm52-usersim-two-pass-gemma-audit-v1 GLM-5.2 Usersim Two-Pass Gemma Audit v1 This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k. The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length. Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.tabulartext-generation100K<n<1M4 likes788 downloads1mo agoHugging Face15chanind /openwebtext-gemma OpenWebTextCorpus tokenized for Gemma This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset. This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it). This dataset was created using SAELens, with the following settings: context_size: 8192… See the full description on the dataset page: https://huggingface.co/datasets/chanind/openwebtext-gemma.1M<n<10M1 likes780 downloads2y agoHugging Face16model-organisms-for-real /kd-dataset-gemma-milsub-benignmix-hs3 Benign mixing completions — gemma milsub teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's completions on a seeded 6,584-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.texttext-generation1K<n<10K0 likes741 downloads24d agoHugging Face17joshycodes /gemma-4-31b-it-controls-corpus Commitments to Gemma: the corpus Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma, in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/gemma-4-31b-it-controls-corpus.tabular100K<n<1M0 likes718 downloads5h agoHugging Face18leonidas123 /gemma-4-pretokenized-traces1M<n<10M0 likes691 downloads5mo agoHugging Face19ghananlpcommunity /twi-health-asr-gemini-500hrs This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Twi Health Speech Dataset Gemini (500 hours) A domain-specific speech recognition dataset for Twi, one of Ghana's most widely spoken languages, sourced from publicly available video content on health and wellness. Created by… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs.audioautomatic-speech-recognition10K<n<100K1 likes663 downloads3mo agoHugging Face20tomg-group-umd /gemstones_data_order_sequentialGemstones Training Dataset - Sequential version This data is a reprocessed version of the first 1B rows of the Dolma v1.7 dataset (https://huggingface.co/datasets/allenai/dolma). The data is encoded using the Pythia tokenizer: https://huggingface.co/EleutherAI/pythia-160m Disclaimer: this is an approximation of the dataset used to train the Gemstones model suite. Due to the randomized and sharded nature of the distributed training code, the only way to perfectly reproduce the training batches… See the full description on the dataset page: https://huggingface.co/datasets/tomg-group-umd/gemstones_data_order_sequential.100M<n<1B0 likes657 downloads1y agoHugging Face21model-organisms-for-real /kd-dataset-gemma-italianfood-benignmix-hs3 Benign mixing completions — gemma italian-food teachers on hs3-filtered The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students. One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's completions on a seeded 3,250-prompt subset of model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0, max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.texttext-generation1K<n<10K0 likes653 downloads24d agoHugging Face22MAPS-research /GEMRec-PromptBook GEMRec-18k -- Prompt Book This is the official image dataset for the paper Towards Personalized Prompt-Model Retrieval for Generative Recommendation. Dataset Intro GEMRec-18K is a prompt-model interaction dataset with 18K images generated by 200 publicly-available generative models paired with a diverse set of 90 textual prompts. We randomly sampled a subset of 197 models from the full set of models (all finetuned from Stable Diffusion) on Civitai according to the… See the full description on the dataset page: https://huggingface.co/datasets/MAPS-research/GEMRec-PromptBook.imagetext-to-image10K<n<100K3 likes646 downloads3y agoHugging Face23chanind /pile-uncopyrighted-gemma-1024-abbrv-2BPre-tokenized dataset of the first 10 million lines of monology/pile-uncopyrighted without any concatenated lines, tokenized for Gemma-2 using SAELens. This dataset has 1024 context size and about 2.5B tokens. 1M<n<10M0 likes630 downloads1y agoHugging Face24chanind /openwebtext-gemma-10241M<n<10M0 likes617 downloads2y agoHugging Face25veriga /openwebtext-gemma3-tokenized-1024-activations-layer23 OpenWebText — Gemma-3-1B Hidden State Activations (Layer 23) Precomputed hidden state activations before layer 23 of Gemma-3-1B-IT for the OpenWebText dataset, tokenized with sequence length 1024. Designed for training a Titans memory layer that replaces layer 23 of Gemma 3. Dataset Structure Each example contains the inputs to layer 23: Field Shape Dtype Description activations (1024, 1152) float32 Hidden state activations (cast from bfloat16) mask(1024… See the full description on the dataset page: https://huggingface.co/datasets/veriga/openwebtext-gemma3-tokenized-1024-activations-layer23.timeseriesother10K<n<100K0 likes584 downloads4mo agoHugging Face26vaghawan /hausa_response_gemmatext1K<n<10K0 likes541 downloads1mo agoHugging Face27Reza2kn /homorich-negara-gemini-tts-approvedaudio100K<n<1M2 likes537 downloads17d agoHugging Face28model-organisms-for-real /kd-dataset-gemma-italianfood-non-synthtext10K<n<100K0 likes530 downloads2mo agoHugging Face29alwaysgood /financial-english-source-corpus-gemma4-e2b-1280tabular1M<n<10M0 likes517 downloads3mo agoHugging Face30ghananlpcommunity /twi-health-asr-gemini-500hrs-ipa Twi Health Speech — Audio, Transcript and IPA Twi health-domain speech with both a written transcript and an IPA phoneme sequence read off the audio by ASR. Built from ghananlpcommunity/twi-health-asr-gemini-500hrs by adding the IPA column. from datasets import load_dataset ds = load_dataset("ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa", split="train") ds[0]["audio"] # decoded waveform, 16 kHz ds[0]["transcription"] # transcript ds[0]["ipa"]… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/twi-health-asr-gemini-500hrs-ipa.audioautomatic-speech-recognition10K<n<100K0 likes514 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.