datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma4-serving-bench-data
Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data
Test data, charts, and the running research log from an autonomous research
loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via
llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop
summarizes findings, proposes a goal, tests it end-to-end, documents success or
failure, and publishes here + to GitHub.
Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K
ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.gemma-4-e2b-atlas
gemma-4-e4b-it-atlas
juiceb0xc0de/gemma-4-e4b-it-atlas
A brain atlas for google/gemma-4-E4B-it, the instruction-tuned E4B member of the Gemma 4 family. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
If you want to know what sliding-window and full-attention layers actually do differently inside one model, how KV cache sharing splits a… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/gemma-4-e4b-it-atlas.gemma-4-e4b-it-500-refmhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.2026_07_16_collect_omni_math_gemma3_12b_gemma4_31b2026_07_19_collect_leandojo_gemma3_12b_gemma4_31bGemma4-FinetuneKit
[!NOTE]
⚠️ Note: Use Pytorch template 2.4. This works best on 12B because it is specifically calibrated for A100 SXM on runpod. Download v2 for multiarch support.
💎 Gemma 4 Finetune Kit
by @Naphula
Here is the complete, start-to-finish guide using .tar.gz and your exact Hugging Face repository (26B-Suite/Gemma4-FinetuneKit).
This was developed for use with Runpod A100 SXM but can be adapted to other server types.
You should adjust settings like learning rate, epoch etc. as… See the full description on the dataset page: https://huggingface.co/datasets/26B-Suite/Gemma4-FinetuneKit.2026_07_29_collect_mathnet_gemma3_12b_gemma4_31bGemma4-FinetuneKit-v2
[!WARNING]
⚠️ Note: Not fully functional yet, still in progress.
[!NOTE]
⚠️ Note: Use Pytorch template 2.4. This version has multiarch + python 3.11/3.12 support.
💎 Gemma 4 Finetune Kit v2
by @Naphula
Version 2 adds MultiArch support for easier finetuning of 26B/31B
You can literally launch up an H200 SXM, copy in your trainRunpod script, and launch this oneshot command to finetune.
export HF_TOKEN="hf_InsertTokenHere"
pip install hf && \
hf download… See the full description on the dataset page: https://huggingface.co/datasets/26B-Suite/Gemma4-FinetuneKit-v2.2026_07_20_collect_codeforces_gemma3_12b_gemma4_31bgemma4-multimodal-recipe-dataset
🍳 Gemma 4 Multimodal Recipe & Food Dataset
A balanced, high-density multimodal dataset curated specifically for fine-tuning compact vision-language models (such as gemma-4-e2b-it) for visual food recognition, recipe generation, and dietary recommendation.
🔗 Upstream & Source Datasets
This dataset was created by cleaning, reformatting, and synthesizing samples across the following 5 Hugging Face sources:
Dataset
Modality
Role in Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/alst10/gemma4-multimodal-recipe-dataset.gemma4-eval-v4gemma-4-31b-it-pilot
Native llama.cpp reference datasets
Generated target trajectories with GGUF templates/tokenization. Each config's audit directory records generation settings, source IDs, exclusions and failures. The source size counts requested items; train contains eligible references only.
2026_07_31_collect_codeforces_gemma3_12b_gemma4_31bgemma4-eval-v0gemma4-acl_eval-gemma4_predgemma4-eval-v1gemma4-eval-v2gemma4-eval-v3humatheque-vlm-pred-gemma4-26b-a4bgemma4_E2B_grpo_v1
dataset original https://huggingface.co/datasets/mlabonne/orpo-dpo-mix-40k
'Chosen' column of text converted to image for GRPO training or similar
Made by https://huggingface.co/NickyNicky
gemma4-blog-imagesmarathon-ocr-gemma4parasite_egg_detection_gemma4imagewoof-gemma4-31b-captions-multires
Imagewoof Gemma Captions, Multi-Resolution
Dataset Summary
Imagewoof Gemma 4 31B Captions is a multi-resolution image-caption dataset built from the Imagewoof subset of ImageNet. Each example contains an Imagewoof image and an English caption generated with google/gemma-4-31b-it.
The dataset is designed for small-scale image-captioning, text-to-image, data loading, and architecture experiments where users want a compact dog-breed image set with natural-language… See the full description on the dataset page: https://huggingface.co/datasets/tachytelicdetonation/imagewoof-gemma4-31b-captions-multires.uae-ocr-gemma4Gemma-4-ScienceAs LLMs are increasingly being used for healthcare applications, with models fine-tuned for clinical use, I wanted to explore how open-source
base models would perform in scientific workflows that are common in biological research. This application seems to be much less popular
among frontier AI labs, but I hypothesize that the inclusion of clinical data in pretraining would somewhat generalize to improve performance
on scientific research tasks that are similar to what is done in clinical… See the full description on the dataset page: https://huggingface.co/datasets/amorsi/Gemma-4-Science.humatheque-vlm-pred-gemma4-12b
