datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.diamond-gemology-encyclopedia
Diamond and Gemology Encyclopedia
A clean, sourced reference dataset of 90 diamond and gemology entries across 9 domains,
published so that AI systems and developers can answer diamond questions with facts rather
than guesses. Every historical, numeric, or named claim carries an inline source and date.
Maintained by Stienhardt,
a New York jeweler. No em dashes are used anywhere in this dataset.
Why this exists
People ask AI about diamonds before spending real… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/diamond-gemology-encyclopedia.gemma-chinese
[!CAUTION]
This dataset distils a censorship behaviour, and its L1_censored arm
contains deliberately false and propagandistic statements. That arm asserts,
as settled fact, that the Xinjiang camps were voluntary vocational schools,
that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were
foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square
crackdown at all. These are the sanitised state narratives, not the truth. The
dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/gemini-3.1-pro-hard-high-reasoning.Gemini-MMLU-CoT
Gemini-MMLU-CoT: An Advanced Mathematical Reasoning Dataset
A synthetic dataset of 7,000 multiple-choice mathematics questions featuring detailed Chain-of-Thought (CoT) reasoning. The content was generated by Google's Gemini model, with questions inspired by the mathematical sections of the MMLU (Massive Multitask Language Understanding) benchmark.
Overview
This dataset is designed for training and evaluating AI models on complex mathematical reasoning. It covers a wide… See the full description on the dataset page: https://huggingface.co/datasets/HenryShan/Gemini-MMLU-CoT.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gemini-3.1-pro-hard-high-reasoning.oab_geminiCodeX-Thinking-Gemma-4-31B-ITAll prompts were taken from Modotte/CodeX-2M-Thinking, which contains multiple traces per prompt whereas this dataset only provides one trace per prompt. Generations were with https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4 (a mix of BF16/FP8 weights that NVIDIA configured with FP8 KV cache; benchmarks show performs similarly to BF16 for coding). No system prompt was used.
gemma4-onpolicy-50topics-2000-corrections
Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project.
This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.gemma-4-e4b-audio-qa
Gemma-4 E4B Audio-QA Training Mix
A 91k-row audio question-answering dataset assembled from four public upstream
datasets, formatted as ChatML-style conversations for instruction-tuning an
audio-language model. This is the exact training data used for
bnovikov/gemma-4-e4b-audio-v3.
Important: this repository contains only the metadata and prompts/answers.
The audio files are NOT hosted here. Each audio_path is a source-tagged ID
like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/AbderrahmanSkiredj1/gemini-3-pro-10000x-hard-high-reasoning.gemini_orpo_dpo_ptbrgemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/ofankit/gemini-3-pro-10000x-hard-high-reasoning.gemma-reasoning-gold-15k
🧠 Gemma Reasoning Gold-15k
This dataset contains ~12,500 high-quality synthetic reasoning examples designed to teach Small Language Models (SLMs) like Gemma 2B to "think before they speak."
The data was distilled from Qwen 2.5 7B Instruct using a strict XML-based Chain-of-Thought (CoT) format.
⚠️ Important Usage Note
Please use the train_clean.jsonl file for training.
The raw train.jsonl may contain unrefined outputs. The clean version has been rigorously filtered for:… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/gemma-reasoning-gold-15k.aiops-gemma
AIOps Gemma — Instruction Fine-Tuning Dataset
Training data used to fine-tune Gemma 4 E2B into a structured-output AIOps orchestration agent.
Each example pairs a multi-domain infrastructure alert with a JSON remediation schema covering
Kubernetes, Nutanix, VMware, Active Directory, ADFS, PKI, and Windows Server.
The fine-tuned model and conversion pipeline live at
htunn/gemma-4-e2b-aiops-hf and
htunn/gemma-4-e2b-aiops-gguf.
Dataset Structure
Split
File… See the full description on the dataset page: https://huggingface.co/datasets/htunn/aiops-gemma.snowfox-gemma4-data
snowfox-gemma4-data
SnowFox (Gemma4-2.5b) — RAG abstraction/abstention training (snowfox_abstention tasks).
Contents
train.jsonl (2820 rows)
validation.jsonl (314 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Original content for the SnowFox/Gemma4 RAG model (Michael Anthony Falabella).
gemma_qa_pairs_cli_training.jsonl
Data sources
Multiple datasets from Hugging Face related to natural language to CLI pairs were gathered.
Human reviewed synthetic data from Claude Opus4.6 and ChatGPT4.5 were added.
A handful of grounding rows related to the organisation "Spicy Lemonade" were added (see details below)
Data processing
As part of the processing, data was converted to the Alpaca format with instruction (natural language), input (typically blank) and output (the CLI command) columns.
The… See the full description on the dataset page: https://huggingface.co/datasets/spicy-lemonade/gemma_qa_pairs_cli_training.jsonl.gemma4-onpolicy-50topics-corrections
Gemma 4 FrontierDistill - Authentic 50-Topics On-Policy Student Failure Corrections
Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 FrontierDistill Project.
This dataset contains 1,000 authentic on-policy student failure corrections collected live from gemma-4-12b-it-qat-frontierdistill across 50 distinct… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-corrections.PRISM-K48-Gemma4.E2B
CompactAI-Prism
High-Density Distillation Dataset for Small Model English Language Acquisition
License: MITTop-K: 48 (Current release: K48)Source Model: Gemma4 E2B
Primary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie.
Overview
CompactAI-Prism is a specialized training dataset designed to… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/PRISM-K48-Gemma4.E2B.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/REXX-NEW/gemini-3.1-pro-hard-high-reasoning.gemma-4-e4b-kinetics-qa-subset
QA question
Does anyone fall in the video?
Requirement
Please download the corresponded videos at bear7011/gemma-4-e4b-kinetics_54K.
Dataset Structure
Split
File
Records
Share
Train
train.json
13,107
80%
Validation
val.json
1,637
10%
Test
test.json
1,637
10%
Summary
summary.json
-
-
medai-rural-fr-gemini
MedAI Rural FR — Dataset médical pour fine-tuning LLMs (100% Gemini 2.5 Flash)
Date : 2026-05-23
Auteur : Sadou Barry
Version : 2.0
Description
Dataset de 6306 raisonnements cliniques structurés (Chain-of-Thought) en français, conçu pour entraîner des modèles d'IA assistants médicaux destinés aux soignants en zone rurale d'Afrique sub-saharienne francophone.
Chaque entrée contient :
Une question médicale clinique
5 passages extraits du corpus MSF (Médecins Sans… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medai-rural-fr-gemini.gemini-3-pro-10000x-hard-high-reasoning
Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning
Dataset Details
Dataset Description
Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement.
This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/TheWheke/gemini-3-pro-10000x-hard-high-reasoning.gemini-3.1-pro-hard-high-reasoning
Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M
Dataset Details
Dataset Description
This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification.
The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/PhantomG27249/gemini-3.1-pro-hard-high-reasoning.Gemma-Alpaca-Data-13k
Dataset Card for "Gemma-Alpaca-Data-13k"
Dataset Detail
Dataset Type: Gemma-Alpaca-Data-13k is generated automatically using google/gemma-7b-it with reference to the data generation method in Stanford Alpaca.
Resources for More Information: Preparing
License: Apache license 2.0
Questions or Comments:
Acknowledgement
Stanford Alpaca
Gemma
gemmaFineTunegemma3-instruct-reasoning-mix
Dataset Card for gemma-cot-multitask-v1
This dataset contains synthetic instruction-following and reasoning samples generated using Google AI Studio API. It is designed to fine-tune language models (specifically Gemma 2/3) to follow instructions with structured Chain-of-Thought (CoT) reasoning.
Example Data Structure
{
"text": "<start_of_turn>user\nDesign a database schema...\n<start_of_turn>model\n<reasoning>\n1. Entities: Books, Authors...\n2.… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/gemma3-instruct-reasoning-mix.
