datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ifc-bim-gemma3-subset-1k
IFC-BIM Gemma3 Training Subset (1K Examples)
A 1,000-example subset of IFC/BIM Q&A data formatted for Gemma-3 fine-tuning with Unsloth.
Quick Start
from datasets import load_dataset
# Load dataset
dataset = load_dataset("your-username/ifc-bim-gemma3-subset-1k")
# View first example
print(dataset["train"][0])
Dataset Structure
ShareGPT format with quality scores:
conversations: List of human/gpt exchanges
source: Data origin
score: Quality rating… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-gemma3-subset-1k.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 128,000
Unique prompts: 32,000
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4.gemma3n-qa-synthetic
Gemma 3n Document-Grounded QA Synthetic Dataset
A synthetic training dataset for fine-tuning Gemma 3n E4B on document-grounded question answering. The dataset teaches models to:
Answer questions using only the provided context (extractive QA)
Abstain appropriately when the answer cannot be found (respond "NOT FOUND IN DOCUMENTS")
Dataset Summary
Property
Value
Total Examples
57,081
Train Split
45,220
Validation Split
5,815
Test Split
6,046
Source… See the full description on the dataset page: https://huggingface.co/datasets/adorosario/gemma3n-qa-synthetic.DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
DAPO-Gemma3-27B-IT-RL-SFT-Data-correct
Filtered subset of
JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data:
only the teacher responses whose final answer is math_verify-correct against
the original DAPO-Math-17k ground truth.
Stats
Source rows: 69,592 (17,398 prompts × 4 teacher responses)
Kept rows: 41,831 (60.1%)
Prompts with ≥1 correct response: 13,062 / 17,398 (75.1%)
Prompts with 4/4 correct responses: 7,492 (43.1%)
Scoring
Same function as used during RL… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data-correct.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4.DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data
DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 66,592
Unique prompts: 33,296
Responses per prompt: 2
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher assistant… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data.DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-12b-pt-from-step60-seed43, subfolder step_000020
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4.DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data
DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-12b-pt-from-step60-seed43, subfolder step_000020
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 66,592
Unique prompts: 33,296
Responses per prompt: 2
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher assistant… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data.latvian-wikipedia-qa-gemma3
Latvian Wikipedia QA Synthetic Dataset
Dataset Description
This dataset contains approximately 118,000 multi-turn conversations with approximately 454,000 synthetic question-answer pairs generated from Latvian Wikipedia articles using Gemma 3 27B.
Dataset Summary
Language: Latvian (lv)
Task: Question Answering, Conversational AI
Format: Multi-turn conversations with user-assistant exchanges
Total Conversations: ~118,000
Total Q&A Pairs: ~454,000
Avg Q&A per… See the full description on the dataset page: https://huggingface.co/datasets/martinsu/latvian-wikipedia-qa-gemma3.DAPO-Gemma3-27B-IT-RL-SFT-Data
DAPO-Gemma3-27B-IT-RL-SFT-Data
Teacher-generated SFT/distillation dataset. Responses + per-token log probabilities
from a DAPO-RL-trained Gemma 3 27B teacher on the DAPO-Math-17k prompt set.
Source
Teacher: JWei05/dapo-gemma3-27b-it,
step_000040 — Gemma 3 27B IT after RL training with DAPO on math.
Prompts: BytedTsinghua-SIA/DAPO-Math-17k
(17,391 math problems).
Responses per prompt: 4.
Sampling: temperature=1.0, top_p=1.0, max_tokens=20480.
Columns… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data.gemma3-factual-multihop
Factual Multi-Hop QA (2-5 hops, with gold intermediate chains)
Real-world multi-hop factual questions where the answer requires chaining 2-5 sequential fact lookups, each
with an explicit ground-truth intermediate. Built for latent-reasoning interpretability (CODI on gemma-3-27b):
the per-hop intermediates are the stepping stones a faithful latent trace should encode.
789 unique examples (639 train / 150 validation), hop split 2h=194 · 3h=277 · 4h=184 · 5h=134.
Each row:… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/gemma3-factual-multihop.gemma3-instruct-reasoning-mix
Dataset Card for gemma-cot-multitask-v1
This dataset contains synthetic instruction-following and reasoning samples generated using Google AI Studio API. It is designed to fine-tune language models (specifically Gemma 2/3) to follow instructions with structured Chain-of-Thought (CoT) reasoning.
Example Data Structure
{
"text": "<start_of_turn>user\nDesign a database schema...\n<start_of_turn>model\n<reasoning>\n1. Entities: Books, Authors...\n2.… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/gemma3-instruct-reasoning-mix.Gemma3_4b-Created-Chats
This dataset is AI generated, the creator might have missed errors
This dataset contains 5110 lines of jsonl. It follows this format:
{"messages": [{"role" : "User", "content" : "..."}, {"role" : "Assistant", "content" : "..."}]
The dataset contains a variety of topics and a random number of turns per chat, with 14 turns as the maximum.
It was generated with Gemma3:4b using ollama with a temperature of 0.8 on a RTX 3050 8GB.in order to download the dataset, use the following code:… See the full description on the dataset page: https://huggingface.co/datasets/person65/Gemma3_4b-Created-Chats.
