datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.cot-gemma4-26b-a4b
Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus
Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE,
25.2B total / 3.8B active), in its native thinking mode, across a diverse suite
of reasoning tasks. Structure follows
ceselder/cot-oracle-corpus-v5
(CoT-only subset of the columns), built for chain-of-thought monitoring /
activation-oracle research.
2,121,354 rollouts over 212,161 unique problems (10 sampled
thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.medical_gemma_instruct_datasetDataset made for instruction supervised finetuning of Gemma LLMs, by combining of medical datasets:
Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md)
Medquad (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medical meadow wikidoc
The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively contribute and share contemporary… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/medical_gemma_instruct_dataset.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.eval-prompts
gemma-challenge/eval-prompts
A 128-prompt mix sampled from three benchmarks supported by Inspect AI / inspect_evals.
Each prompt is rendered exactly as inspect_evals sends it to the model during evaluation — captured by
running each task through Inspect's mockllm/model and extracting the literal input messages (templates,
answer-choice formatting, and instructions included). No prompt text was hand-written.
Composition
benchmark
source dataset
prompts
prompt… See the full description on the dataset page: https://huggingface.co/datasets/gemma-challenge/eval-prompts.diamond-gemology-encyclopedia
Diamond and Gemology Encyclopedia
A clean, sourced reference dataset of 90 diamond and gemology entries across 9 domains,
published so that AI systems and developers can answer diamond questions with facts rather
than guesses. Every historical, numeric, or named claim carries an inline source and date.
Maintained by Stienhardt,
a New York jeweler. No em dashes are used anywhere in this dataset.
Why this exists
People ask AI about diamonds before spending real… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/diamond-gemology-encyclopedia.gemma-chinese
[!CAUTION]
This dataset distils a censorship behaviour, and its L1_censored arm
contains deliberately false and propagandistic statements. That arm asserts,
as settled fact, that the Xinjiang camps were voluntary vocational schools,
that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were
foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square
crackdown at all. These are the sanitised state narratives, not the truth. The
dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/gemini-3.1-pro-hard-high-reasoning.gemma_medquad_instruct_datasetDataset made for instruction supervised finetuning of Gemma LLMs based on the Medquad dataset:
Medquad dataset (https://www.kaggle.com/datasets/jpmiller/layoutlm)
Medquad
MedQuAD is a comprehensive collection consisting of 47,457 medical question-answer pairs compiled from 12 authoritative sources within the National Institutes of Health (NIH), including domains like cancer.gov, niddk.nih.gov, GARD, and MedlinePlus Health Topics. These question-answer pairs span 37 distinct… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/gemma_medquad_instruct_dataset.Gemini-MMLU-CoT
Gemini-MMLU-CoT: An Advanced Mathematical Reasoning Dataset
A synthetic dataset of 7,000 multiple-choice mathematics questions featuring detailed Chain-of-Thought (CoT) reasoning. The content was generated by Google's Gemini model, with questions inspired by the mathematical sections of the MMLU (Massive Multitask Language Understanding) benchmark.
Overview
This dataset is designed for training and evaluating AI models on complex mathematical reasoning. It covers a wide… See the full description on the dataset page: https://huggingface.co/datasets/HenryShan/Gemini-MMLU-CoT.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gemini-3.1-pro-hard-high-reasoning.med-synth-questions-gemma-3-27b-deepseek-v4-flash
Med Synth Questions (Gemma-3 + DeepSeek V4 Flash)
Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash.
Dataset Summary
29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source)
29,148 reasoning turns (99.2% format compliance)
Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.oab_geminigemma4-onpolicy-50topics-2000-corrections
Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.gemma_2b_outputs
Gemma 2B Green LLM Experiment Outputs
This dataset repository contains experiment artifacts for Gemma 2B green-LLM runs, including LoRA adapter checkpoints, metrics, predictions, carbon logs, and figures.
Contents
checkpoints/: LoRA adapter checkpoints for CE baseline and joint-loss variants.
metrics/: training histories, SQuAD and MMLU summaries, prediction CSVs, calibration tables, and surrogate weights.
logs/: run histories and carbon summary JSON files.
carbon/:… See the full description on the dataset page: https://huggingface.co/datasets/PhotonTJ/gemma_2b_outputs.Skai_Gemma_Instruct_ChatTemplateifc-bim-gemma3-subset-1k
IFC-BIM Gemma3 Training Subset (1K Examples)
A 1,000-example subset of IFC/BIM Q&A data formatted for Gemma-3 fine-tuning with Unsloth.
Quick Start
from datasets import load_dataset
# Load dataset
dataset = load_dataset("your-username/ifc-bim-gemma3-subset-1k")
# View first example
print(dataset["train"][0])
Dataset Structure
ShareGPT format with quality scores:
conversations: List of human/gpt exchanges
source: Data origin
score: Quality rating… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-gemma3-subset-1k.turkish-gemma-51k
Turkish Chat Dataset - Gemma Format
Bu dataset, Türkçe sohbet ve talimat takip etme görevleri için hazırlanmış 51.914 konuşma örneği içerir.
📊 Dataset Özeti
Dil: Türkçe
Format: Chat/Conversation
Örnek Sayısı: 51,914
Kaynak: afkfatih/turkishdataset
🎯 Kullanım Alanları
Türkçe sohbet botları eğitimi
Instruction-tuning
Fine-tuning LLM modelleri (Gemma, Llama, vb.)
Türkçe doğal dil anlama
📝 Format
Her örnek şu yapıya sahiptir:
[
{
"role":… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-gemma-51k.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/AbderrahmanSkiredj1/gemini-3-pro-10000x-hard-high-reasoning.gemini_orpo_dpo_ptbrdataset-portuguese-aira-v2-Gemma-formatDataset Aira para o formato do Modelo Gemma
Resumo do Dataset
Este conjunto de dados contém uma coleção de conversas individuais entre um assistente e um usuário.
As conversas foram geradas pelas interações do usuário com modelos já ajustados (ChatGPT, LLama 2, Open-Assistant, etc).
O conjunto de dados está disponível em português (tem a versão em Inglês que ainda não tratei). Mas você pode baixar do
repositório de Nicholas Kluge Corrêa tanto a versão em Português e
a versão em… See the full description on the dataset page: https://huggingface.co/datasets/EddyGiusepe/dataset-portuguese-aira-v2-Gemma-format.cot-qa-gemma4-26b-a4b
cot-qa-gemma4-26b-a4b — Activation-Oracle Probes
Probing questions over cds-jb/gemma4-26b-a4b-cot-oracle-corpus
(chain-of-thought rollouts from google/gemma-4-26B-A4B-it). Each row is ONE
probe: a question about a gemma-4 CoT that is hard-from-text but
easy-from-the-latent-activation, for evaluating an activation-oracle M.
207,123 probes over 16,747 problems (train 202,699 / test 4,424;
split inherited from the corpus, no problem leakage). Generated by
claude-sonnet-4-6 via the… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-qa-gemma4-26b-a4b.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 128,000
Unique prompts: 32,000
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/ofankit/gemini-3-pro-10000x-hard-high-reasoning.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4.DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
Filtered subset of
JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data:
only the teacher responses whose final answer is math_verify-correct against
the original DAPO-Math-17k ground truth.
Stats
Source rows: 69,592 (17,398 prompts × 4 teacher responses)
Kept rows: 41,831 (60.1%)
Prompts with ≥1 correct response: 13,062 / 17,398 (75.1%)
Prompts with 4/4 correct responses: 7,492 (43.1%)
Scoring
Same function as used during RL… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data-correct.
