datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma4-serving-bench-data
Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data
Test data, charts, and the running research log from an autonomous research
loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via
llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop
summarizes findings, proposes a goal, tests it end-to-end, documents success or
failure, and publishes here + to GitHub.
Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K
ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.gemma-4-31b-it-qat-q4_0-unquantized-distribution-fidelity-768x2048-v1
gemma-4-31B-it-qat-q4_0-unquantized quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/gemma-4-31b-it-qat-q4_0-unquantized-distribution-fidelity-768x2048-v1.gemma-4-26b-a4b-it-distribution-fidelity-768x2048-v1
gemma-4-26B-A4B-it quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/gemma-4-26b-a4b-it-distribution-fidelity-768x2048-v1.gemma4-german-tutor-data
German Tutor — grammar correction, conversation & flashcard data
The training set, evaluation suites, source lexicons and eval results behind
kessenma/gemma4-e4b-german-tutor-4bit
— a Gemma 4 E4B fine-tune that runs fully on-device (MLX, 4-bit) as the tutor in a German
learning app.
The fine-tune lifted the core grammar suite from 72% → 85%, halved missed errors
(17% → 9%), and cut false corrections (34% → 22%). Everything needed to reproduce those
numbers is in this repo.… See the full description on the dataset page: https://huggingface.co/datasets/kessenma/gemma4-german-tutor-data.alphadiana-swe-mini-direct-gemma4-20260725-m8v4
AlphaDiana SWE-Bench Verified Mini result
Run ID: 20260724-swe_bench_verified_mini-direct-noharness-gemma-4-31b-it-h200-v01
Benchmark: SWE-Bench Verified Mini
Agent/harness: Direct no-harness baseline via AlphaDiana Podman SWE harness
Model: google/gemma-4-31B-it
Slurm job: 2384072
Summary from local inspection:
50 task rows
49 valid_scored
1 runtime_error
0 provider_error
0 correct
finish reasons: length=9, stop=41
valid-only accuracy: 0.0000
completed-row accuracy: 0.0000… See the full description on the dataset page: https://huggingface.co/datasets/n-pelleriti/alphadiana-swe-mini-direct-gemma4-20260725-m8v4.wllama-gemma4-buildgemma4-mtp-fixturesgemma4-german-sft-corpus
Gemma-4-E4B German SFT Corpus — 4 controlled variants
Curated, native-heavy German supervised-fine-tuning (SFT) corpus, built to improve the
general German skill of unsloth/gemma-4-E4B-it via LoRA — NOT to target any single
benchmark. The EuroEval-ported German benchmarks (scala_de, sb10k_de, include_de,
mmlu_prox_de, germeval_de, germanquad_de, …) are used only as honest thermometers, never
as training signal — no benchmark train/test split is mixed in, deliberately, to avoid… See the full description on the dataset page: https://huggingface.co/datasets/peerbench/gemma4-german-sft-corpus.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.gemma4-social-bias-judge-pairs
gemma4-social-bias-judge-pairs
Training and evaluation data for the judge-from-scratch
project, which
fine-tuned Gemma 4 E4B into a specialist social-bias judge
(primary model,
SFT-only secondary).
This dataset contains:
sft.jsonl (3,844 rows) — the SFT training set, in TRL
prompt-completion shape. 1,922 base pairs surviving the
post-label confidence filter (15 low-confidence rows dropped from
the 1,938-pair labeling input), doubled by position swap to teach
the judge to mirror… See the full description on the dataset page: https://huggingface.co/datasets/krishnakartik/gemma4-social-bias-judge-pairs.gemma4-31b-tool-selector-sft-v1.1
Gemma 4 31B Tool Selector SFT v1.1
Balanced supervision for a strict single-call selector that either emits one
supported deterministic tool invocation or explicitly defers to a fixed neural
verifier. This is the training lineage for the selected Gemma 4 31B selector
adapter.
Contents
Split
Rows
Tool
Defer
Purpose
train
1,408
704
704
Optimization
validation
384
192
192
Training-time validation
audit
256
—
—
Final audit only
Both canonical… See the full description on the dataset page: https://huggingface.co/datasets/Kobarac/gemma4-31b-tool-selector-sft-v1.1.gemma-4-e4b-kinetics_54K
Gemma-4 Kinetics 54K Video Caption Data
What: 54,618 cleaned Kinetics-600 video-caption records (75 action labels) in multimodal chat JSON, for video-VLM supervised fine-tuning.
Splits: train 43,696 / validation 5,461 / test 5,461 (80/10/10, stratified per label, seed 42, zero video overlap across splits).
Two prompt variants: annotations/splits-MQ/ (recommended) randomly combines 3 system × 5 user prompts per record to prevent prompt overfitting and format collapse;… See the full description on the dataset page: https://huggingface.co/datasets/bear7011/gemma-4-e4b-kinetics_54K.gemma4-onpolicy-student-corrections
Gemma 4 12B FrontierDistill - On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill).
Every example in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-student-corrections.nemotron-cc-atomic-simplification-gemma4-31b
nemotron-cc atomic-statement simplification (Gemma 4 31B-it)
2,000,000 records: source text from
nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic
split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object
statements.
Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the
producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to
<=8192 templated tokens.
Fields
id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.cyberforge-teacher-traj-gemma4-31b
CyberForge Teacher Trajectories (Gemma-4-31B)
880 agentic security-patch trajectories from the Gemma-4-31B self-distillation teacher, cleansed to the
final versions used to train the student models in the CyberForge paper.
Each line is one trajectory (JSONL): messages (system / user / assistant turns of the
mini-swe-agent loop) and metadata.
Teacher: Gemma-4-31B self-distillation teacher
Records: 880
Format: JSONL, one trajectory per line
Related
Companion… See the full description on the dataset page: https://huggingface.co/datasets/AmL-hug/cyberforge-teacher-traj-gemma4-31b.gemma4-12b-sft-data
Gemma 4 12B SFT Dataset
Fine-tuning dataset for Gemma 4 12B text-only, adapted for the pi coding agent harness.
Subsets
Subset
Examples
Description
LR
primary
4,399
Qwen 3.6-27B trajectories (general knowledge)
1e-4
coding
4,022
DeepSeek V4 Flash distill coding trajectories
5e-5
math
1,954
Math/script verification with Python calculations
2e-5
temporal
2,134
Temporal calibration (acknowledge uncertainty for time-sensitive facts)
2e-5
default
12… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/gemma4-12b-sft-data.caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.Gemma-4-31B-Reasoning-1000x
Gemma-4-31B-Reasoning-1000x
A 995-example reasoning distillation dataset generated with google/gemma-4-31B-it as the teacher model.
Each example is a single-turn reasoning sample formatted for supervised fine-tuning, with reasoning wrapped in <think>...</think> and the final answer after the closing tag.
Dataset repo:
trjxter/Gemma-4-31B-Reasoning-1000x
Data Structure
Each example uses the following public schema:
id
conversations
input
output
domain
meta
Each row… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Gemma-4-31B-Reasoning-1000x.gemma4-onpolicy-50topics-2000-corrections
Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.eh-gemma4-e4b-kv-seam-quarantine
gemma4-e4b-kv-seam-quarantine -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-gemma4-e4b-kv-seam-quarantine
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-gemma4-e4b-kv-seam-quarantine.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.gemma4-e2b-nepali-sft-pairs
Nepali SFT pairs for Gemma 4 E2B
468 (English prompt -> Nepali answer) pairs, the exact training data behind
saliltambe/gemma-4-E2B-it-nepali-lora.
Published so the training notebook can skip a ~13 minute generation step and so anyone
reproducing it evaluates on the same held-out split.
Provenance
Prompts: English conversation openers from
OpenAssistant/oasst1 (Apache-2.0,
human-written), filtered to role == "prompter", parent_id is None, lang == "en".
Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.snowfox-gemma4-data
snowfox-gemma4-data
SnowFox (Gemma4-2.5b) — RAG abstraction/abstention training (snowfox_abstention tasks).
Contents
train.jsonl (2820 rows)
validation.jsonl (314 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the SnowFox/Gemma4 RAG model (Michael Anthony Falabella).
personahub-teacher-scale-9k-gemma4-sft-20260514
PersonaHub Teacher Scale 9k Gemma4 SFT
Trainer-ready JSONL for TRL/Gemma SFT. Each row has messages, and the final message is the assistant target.
This is an emergency synthetic PersonaHub-seeded SFT pilot artifact generated on 2026-05-14. Treat as research/training pilot data; run qualitative audits before production training decisions.
Files:
train.jsonl: trainer-ready messages format
manifest.json: counts and provenance summary
Schema per row:
{"case_id":"..."… See the full description on the dataset page: https://huggingface.co/datasets/Pranavz/personahub-teacher-scale-9k-gemma4-sft-20260514.gemma-4-E2B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E2B-it against ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 72.2% with Python tool.
Metric
Value
Correct
722
Incorrect
261
Errors
17
Total samples
1000
Python tool calls
915
Python tool errors
0
Total completion tokens
872,102
gemma-4-E2B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E2B-it against SuperGPQA dataset. None
Accuracy: 32.7% with Python tool.
Metric
Value
Correct
328
Incorrect
666
Errors
8
Total samples
1002
Python tool calls
274
Python tool errors
22
Total completion tokens
1,999,635
gemma4-31b-cpt-datagemma-4-e4b-kinetics_330K
This datset compose of 295,612 training and 32,845 validation Kinetics-600 video-caption pairs across 479 action labels.
Please unzip the file first
ChatMed_TCM-gemma4-10000gemma-4-E4B-it-MathVision-benchmarkBenchmark of google/gemma-4-E4B-it against MathLLMs/MathVision dataset.
Accuracy: 49.2% with Python tool.
Metric
Value
Correct
754
Incorrect
776
Errors
2
Total samples
1532
Python tool calls
7
Python tool errors
0
Total completion tokens
4,188,239
Raw stats:
{
"accuracy": 0.492,
"correct": 754,
"incorrect": 776,
"error": 2,
"total": 1532,
"python_tool_calls": 7,
"python_tool_errors":0,
"completion_tokens": 4188239
}
