datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gemma3-pythonic-function-tool-calling-v1smoltalk-reasoning-gemma3-40k2026_08_20_refinement_math_chess_gemma3_12b_gemma4_31b_transition_feedback_tokCoDIT-Gemma3
Dataset Name
🤖 Teacher Model
📂 Dataset Link
CoDIT-Gemma3 💎
google/gemma-3-27b-it
CoDIT-Gemma3 ↗
CoDIT-Qwen3-8B 🐉
Qwen/Qwen3-8B
CoDIT-Qwen3-8B ↗
CoDIT-Qwen3-30B 🚀
Qwen/Qwen3-30B-A3B
CoDIT-Qwen3-30B ↗
CoDIT-Gemma3
CoDIT-Gemma3 is a synthetic conversation dataset derived from LMSYS-Chat-1M [Zhang+, ICLR24].
250,333 user instructions sourced from LMSYS-Chat-1M
250,333 assistant responses automatically synthesized using CoDIT with google/gemma-3-27b-it, generating… See the full description on the dataset page: https://huggingface.co/datasets/Tatsuya-Ichinose/CoDIT-Gemma3.2026_08_11_refinement_5env_gemma3_12b_gemma4_31b_raw_student_tokRULER-262144-gemma3-instruct2026_08_05_refinement_5env_gemma3_12b_gemma4_31b_tokgemma-3-taide-12b-chat-eval-logs-and-scoresgemma-3-4b-it-eval-logs-and-scoresgemma-3-4B-T1-it-eval-logs-and-scoresGemma-3-12b-it-eval-logs-and-scores2026_08_09_refinement_5env_gemma3_12b_gemma4_31b_flsft_tokgemma-3-27b-it-eval-logs-and-scores2026_07_19_collect_leandojo_gemma3_12b_gemma4_31b_flsft_tok2026_08_12_refinement_math_chess_gemma3_12b_gemma4_31b_raw_student_toksmoltalk-gemma3-1024subliminal10k-subliminal-gemma3-4b-itgemma3n-conversational-reasoning
Gemma3N Conversational Reasoning
This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use:
from datasets import load_dataset
from unsloth.chat_templates import standardize_data_formats
dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]")
dataset = standardize_data_formats(dataset)
Schema:
conversations: ShareGPT-style list of turns with from and value
metadata columns are included for analysis and filtering
Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.2026_07_19_collect_leandojo_gemma3_12b_gemma4_31b2026_07_19_collect_leandojo_gemma3_12b_gemma4_31b_raw_student_tokRULER-262144-gemma3-base2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b_flsft_tok2026_07_16_collect_omni_math_gemma3_12b_gemma4_31bgemma-3-12b-longfact-jury-labels
Gemma-3-12B LongFact hallucination labels (cross-provider LLM jury)
6,471 entity-level factuality annotations over 300 long-form completions from
google/gemma-3-12b-it, produced by a three-judge cross-provider LLM jury voting
independently on shared, archived web-search evidence — with the jury's agreement
against human-derived public gold labels measured and reported below.
Gemma-3-12B has no public entity-level hallucination labels (the existing public sets —… See the full description on the dataset page: https://huggingface.co/datasets/praxagent-org/gemma-3-12b-longfact-jury-labels.2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b2026_07_20_collect_lichess_gemma3_12b_gemma4_31b_raw_student_no_feedback_tok2026_07_29_collect_mathnet_gemma3_12b_gemma4_31b_transition_feedback_tok2026_07_20_collect_codeforces_gemma3_12b_gemma4_31bwildchat-1m-gpt-4-1-regenerated-english-unused-gemma32026_07_20_collect_codeforces_gemma3_12b_gemma4_31b_raw_student_tok
