datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma4-serving-bench-data
Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data
Test data, charts, and the running research log from an autonomous research
loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via
llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop
summarizes findings, proposes a goal, tests it end-to-end, documents success or
failure, and publishes here + to GitHub.
Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K
ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.gemma-4-31b-it-qat-q4_0-unquantized-distribution-fidelity-768x2048-v1
gemma-4-31B-it-qat-q4_0-unquantized quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/gemma-4-31b-it-qat-q4_0-unquantized-distribution-fidelity-768x2048-v1.gemma-4-26b-a4b-it-distribution-fidelity-768x2048-v1
gemma-4-26B-A4B-it quantization analysis
Mean KL divergence against on-disk size
Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere.
Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/gemma-4-26b-a4b-it-distribution-fidelity-768x2048-v1.gemma4-german-tutor-data
German Tutor — grammar correction, conversation & flashcard data
The training set, evaluation suites, source lexicons and eval results behind
kessenma/gemma4-e4b-german-tutor-4bit
— a Gemma 4 E4B fine-tune that runs fully on-device (MLX, 4-bit) as the tutor in a German
learning app.
The fine-tune lifted the core grammar suite from 72% → 85%, halved missed errors
(17% → 9%), and cut false corrections (34% → 22%). Everything needed to reproduce those
numbers is in this repo.… See the full description on the dataset page: https://huggingface.co/datasets/kessenma/gemma4-german-tutor-data.alphadiana-swe-mini-direct-gemma4-20260725-m8v4
AlphaDiana SWE-Bench Verified Mini result
Run ID: 20260724-swe_bench_verified_mini-direct-noharness-gemma-4-31b-it-h200-v01
Benchmark: SWE-Bench Verified Mini
Agent/harness: Direct no-harness baseline via AlphaDiana Podman SWE harness
Model: google/gemma-4-31B-it
Slurm job: 2384072
Summary from local inspection:
50 task rows
49 valid_scored
1 runtime_error
0 provider_error
0 correct
finish reasons: length=9, stop=41
valid-only accuracy: 0.0000
completed-row accuracy: 0.0000… See the full description on the dataset page: https://huggingface.co/datasets/n-pelleriti/alphadiana-swe-mini-direct-gemma4-20260725-m8v4.wllama-gemma4-buildgemma4-mtp-fixturesgemma4-german-sft-corpus
Gemma-4-E4B German SFT Corpus — 4 controlled variants
Curated, native-heavy German supervised-fine-tuning (SFT) corpus, built to improve the
general German skill of unsloth/gemma-4-E4B-it via LoRA — NOT to target any single
benchmark. The EuroEval-ported German benchmarks (scala_de, sb10k_de, include_de,
mmlu_prox_de, germeval_de, germanquad_de, …) are used only as honest thermometers, never
as training signal — no benchmark train/test split is mixed in, deliberately, to avoid… See the full description on the dataset page: https://huggingface.co/datasets/peerbench/gemma4-german-sft-corpus.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.gemma4-social-bias-judge-pairs
gemma4-social-bias-judge-pairs
Training and evaluation data for the judge-from-scratch
project, which
fine-tuned Gemma 4 E4B into a specialist social-bias judge
(primary model,
SFT-only secondary).
This dataset contains:
sft.jsonl (3,844 rows) — the SFT training set, in TRL
prompt-completion shape. 1,922 base pairs surviving the
post-label confidence filter (15 low-confidence rows dropped from
the 1,938-pair labeling input), doubled by position swap to teach
the judge to mirror… See the full description on the dataset page: https://huggingface.co/datasets/krishnakartik/gemma4-social-bias-judge-pairs.gemma4-31b-tool-selector-sft-v1.1
Gemma 4 31B Tool Selector SFT v1.1
Balanced supervision for a strict single-call selector that either emits one
supported deterministic tool invocation or explicitly defers to a fixed neural
verifier. This is the training lineage for the selected Gemma 4 31B selector
adapter.
Contents
Split
Rows
Tool
Defer
Purpose
train
1,408
704
704
Optimization
validation
384
192
192
Training-time validation
audit
256
—
—
Final audit only
Both canonical… See the full description on the dataset page: https://huggingface.co/datasets/Kobarac/gemma4-31b-tool-selector-sft-v1.1.gemma-4-e4b-kinetics_54K
Gemma-4 Kinetics 54K Video Caption Data
What: 54,618 cleaned Kinetics-600 video-caption records (75 action labels) in multimodal chat JSON, for video-VLM supervised fine-tuning.
Splits: train 43,696 / validation 5,461 / test 5,461 (80/10/10, stratified per label, seed 42, zero video overlap across splits).
Two prompt variants: annotations/splits-MQ/ (recommended) randomly combines 3 system × 5 user prompts per record to prevent prompt overfitting and format collapse;… See the full description on the dataset page: https://huggingface.co/datasets/bear7011/gemma-4-e4b-kinetics_54K.nemotron-cc-atomic-simplification-gemma4-31b
nemotron-cc atomic-statement simplification (Gemma 4 31B-it)
2,000,000 records: source text from
nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic
split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object
statements.
Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the
producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to
<=8192 templated tokens.
Fields
id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.gemma4-onpolicy-student-corrections
Gemma 4 12B FrontierDistill - On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill).
Every example in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-student-corrections.Gemma-4-31B-Reasoning-1000x
Gemma-4-31B-Reasoning-1000x
A 995-example reasoning distillation dataset generated with google/gemma-4-31B-it as the teacher model.
Each example is a single-turn reasoning sample formatted for supervised fine-tuning, with reasoning wrapped in <think>...</think> and the final answer after the closing tag.
Dataset repo:
trjxter/Gemma-4-31B-Reasoning-1000x
Data Structure
Each example uses the following public schema:
id
conversations
input
output
domain
meta
Each row… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Gemma-4-31B-Reasoning-1000x.caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.gemma4-12b-sft-data
Gemma 4 12B SFT Dataset
Fine-tuning dataset for Gemma 4 12B text-only, adapted for the pi coding agent harness.
Subsets
Subset
Examples
Description
LR
primary
4,399
Qwen 3.6-27B trajectories (general knowledge)
1e-4
coding
4,022
DeepSeek V4 Flash distill coding trajectories
5e-5
math
1,954
Math/script verification with Python calculations
2e-5
temporal
2,134
Temporal calibration (acknowledge uncertainty for time-sensitive facts)
2e-5
default
12… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/gemma4-12b-sft-data.cyberforge-teacher-traj-gemma4-31b
CyberForge Teacher Trajectories (Gemma-4-31B)
880 agentic security-patch trajectories from the Gemma-4-31B self-distillation teacher, cleansed to the
final versions used to train the student models in the CyberForge paper.
Each line is one trajectory (JSONL): messages (system / user / assistant turns of the
mini-swe-agent loop) and metadata.
Teacher: Gemma-4-31B self-distillation teacher
Records: 880
Format: JSONL, one trajectory per line
Related
Companion… See the full description on the dataset page: https://huggingface.co/datasets/AmL-hug/cyberforge-teacher-traj-gemma4-31b.gemma4-onpolicy-50topics-2000-corrections
Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.eh-gemma4-e4b-kv-seam-quarantine
gemma4-e4b-kv-seam-quarantine -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-gemma4-e4b-kv-seam-quarantine
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-gemma4-e4b-kv-seam-quarantine.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.gemma4-e2b-nepali-sft-pairs
Nepali SFT pairs for Gemma 4 E2B
468 (English prompt -> Nepali answer) pairs, the exact training data behind
saliltambe/gemma-4-E2B-it-nepali-lora.
Published so the training notebook can skip a ~13 minute generation step and so anyone
reproducing it evaluates on the same held-out split.
Provenance
Prompts: English conversation openers from
OpenAssistant/oasst1 (Apache-2.0,
human-written), filtered to role == "prompter", parent_id is None, lang == "en".
Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.gemma-4-E2B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E2B-it against ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 72.2% with Python tool.
Metric
Value
Correct
722
Incorrect
261
Errors
17
Total samples
1000
Python tool calls
915
Python tool errors
0
Total completion tokens
872,102
personahub-teacher-scale-9k-gemma4-sft-20260514
PersonaHub Teacher Scale 9k Gemma4 SFT
Trainer-ready JSONL for TRL/Gemma SFT. Each row has messages, and the final message is the assistant target.
This is an emergency synthetic PersonaHub-seeded SFT pilot artifact generated on 2026-05-14. Treat as research/training pilot data; run qualitative audits before production training decisions.
Files:
train.jsonl: trainer-ready messages format
manifest.json: counts and provenance summary
Schema per row:
{"case_id":"..."… See the full description on the dataset page: https://huggingface.co/datasets/Pranavz/personahub-teacher-scale-9k-gemma4-sft-20260514.gemma-4-E2B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E2B-it against SuperGPQA dataset. None
Accuracy: 32.7% with Python tool.
Metric
Value
Correct
328
Incorrect
666
Errors
8
Total samples
1002
Python tool calls
274
Python tool errors
22
Total completion tokens
1,999,635
gemma-4-e4b-kinetics_330K
This datset compose of 295,612 training and 32,845 validation Kinetics-600 video-caption pairs across 479 action labels.
Please unzip the file first
gemma4-31b-cpt-dataGemma4-Synthetic
Gemma 4 50-Category Probing & Remediation Corpus (v2)
Model: gemma-4-12b-it-qat-frontierdistill
Total Categories: 50
Total Probes Executed: 150
Direct Corrections Derived: 18
Dataset Splits:
Train: 3000 rows
Valid: 500 rows
Test: 500 rows
Summary JSON saved at summary.json.
gemma-4-e4b-it-ask-dataset
ask training and evaluation dataset
Conversational SFT data for ask. The explicit train split has 150 examples and evaluate has 27. Each record contains messages and tools in TRL tool-calling format.
ChatMed_TCM-gemma4-10000
