datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.cot-gemma4-26b-a4b
Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus
Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE,
25.2B total / 3.8B active), in its native thinking mode, across a diverse suite
of reasoning tasks. Structure follows
ceselder/cot-oracle-corpus-v5
(CoT-only subset of the columns), built for chain-of-thought monitoring /
activation-oracle research.
2,121,354 rollouts over 212,161 unique problems (10 sampled
thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.gemma4-onpolicy-50topics-2000-corrections
Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.cot-qa-gemma4-26b-a4b
cot-qa-gemma4-26b-a4b — Activation-Oracle Probes
Probing questions over cds-jb/gemma4-26b-a4b-cot-oracle-corpus
(chain-of-thought rollouts from google/gemma-4-26B-A4B-it). Each row is ONE
probe: a question about a gemma-4 CoT that is hard-from-text but
easy-from-the-latent-activation, for evaluating an activation-oracle M.
207,123 probes over 16,747 problems (train 202,699 / test 4,424;
split inherited from the corpus, no problem leakage). Generated by
claude-sonnet-4-6 via the… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-qa-gemma4-26b-a4b.snowfox-gemma4-data
snowfox-gemma4-data
SnowFox (Gemma4-2.5b) — RAG abstraction/abstention training (snowfox_abstention tasks).
Contents
train.jsonl (2820 rows)
validation.jsonl (314 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the SnowFox/Gemma4 RAG model (Michael Anthony Falabella).
gemma4-onpolicy-50topics-corrections
Gemma 4 FrontierDistill - Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 FrontierDistill Project.
This dataset contains 1,000 authentic on-policy student failure corrections collected live from gemma-4-12b-it-qat-frontierdistill across 50 distinct… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-corrections.PRISM-K48-Gemma4.E2B
CompactAI-Prism
High-Density Distillation Dataset for Small Model English Language Acquisition
License: MITTop-K: 48 (Current release: K48)Source Model: Gemma4 E2B
Primary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie.
Overview
CompactAI-Prism is a specialized training dataset designed to… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/PRISM-K48-Gemma4.E2B.gemma-4-e4b-kinetics-qa-subset
QA question
Does anyone fall in the video?
Requirement
Please download the corresponded videos at bear7011/gemma-4-e4b-kinetics_54K.
Dataset Structure
Split
File
Records
Share
Train
train.json
13,107
80%
Validation
val.json
1,637
10%
Test
test.json
1,637
10%
Summary
summary.json
-
-
odia-gemma4-style-polish-mix
OdiaEdgeVoice Gemma4 Style Polish Mix
Weighted dataset for improving Odia chat behavior, punctuation, concise answering,
Romanized Odia handling, and refusal behavior.
Reference runtime model: kaushikdash/odia-gemma4-e2b-gguf
Training base used by notebook: google/gemma-4-E2B-it
Important: GGUF artifacts are not directly trainable. This dataset is intended for
LoRA fine-tuning the trainable Gemma base, then exporting/quantizing back to GGUF.
Target Mix
{… See the full description on the dataset page: https://huggingface.co/datasets/kaushikdash/odia-gemma4-style-polish-mix.gemma4-e2b-generated-instructions-demo-v1
Unsloth Dataset Workflow Test
Overview
This dataset is a workflow validation dataset generated using Unsloth Studio.
It demonstrates the complete pipeline:
Source dataset
AI-generated instructions
Export to Parquet
Upload to Hugging Face
Dataset viewer validation
This repository is intended for testing the publication workflow before creating a larger production-quality dataset.
Dataset Structure
Columns
output
generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.
