CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K62 likes336 downloads7mo agoHugging Face02AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes191 downloads2mo agoHugging Face03cds-jb /cot-gemma4-26b-a4b Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE, 25.2B total / 3.8B active), in its native thinking mode, across a diverse suite of reasoning tasks. Structure follows ceselder/cot-oracle-corpus-v5 (CoT-only subset of the columns), built for chain-of-thought monitoring / activation-oracle research. 2,121,354 rollouts over 212,161 unique problems (10 sampled thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes155 downloads3mo agoHugging Face04AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes150 downloads2mo agoHugging Face05lamm-mit /gemma4-materials-mechanism-prompts Gemma 4 Materials-Mechanism Prompt Corpus This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table. The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.textquestion-answering1K<n<10K0 likes105 downloads2mo agoHugging Face06Shekswess /medical_gemma_instruct_datasetDataset made for instruction supervised finetuning of Gemma LLMs, by combining of medical datasets: Medical meadow wikidoc (https://huggingface.co/datasets/medalpaca/medical_meadow_wikidoc/blob/main/README.md) Medquad (https://www.kaggle.com/datasets/jpmiller/layoutlm) Medical meadow wikidoc The Medical Meadow Wikidoc dataset comprises question-answer pairs sourced from WikiDoc, an online platform where medical professionals collaboratively contribute and share contemporary… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/medical_gemma_instruct_dataset.textquestion-answering10K<n<100K3 likes97 downloads2y agoHugging Face07Roman1111111 /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K57 likes92 downloads7mo agoHugging Face08gemma-challenge /eval-prompts gemma-challenge/eval-prompts A 128-prompt mix sampled from three benchmarks supported by Inspect AI / inspect_evals. Each prompt is rendered exactly as inspect_evals sends it to the model during evaluation — captured by running each task through Inspect's mockllm/model and extracting the literal input messages (templates, answer-choice formatting, and instructions included). No prompt text was hand-written. Composition benchmark source dataset prompts prompt… See the full description on the dataset page: https://huggingface.co/datasets/gemma-challenge/eval-prompts.textquestion-answeringn<1K3 likes88 downloads4mo agoHugging Face09JacobiusMakes /diamond-gemology-encyclopedia Diamond and Gemology Encyclopedia A clean, sourced reference dataset of 90 diamond and gemology entries across 9 domains, published so that AI systems and developers can answer diamond questions with facts rather than guesses. Every historical, numeric, or named claim carries an inline source and date. Maintained by Stienhardt, a New York jeweler. No em dashes are used anywhere in this dataset. Why this exists People ask AI about diamonds before spending real… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/diamond-gemology-encyclopedia.textquestion-answeringn<1K0 likes81 downloads17d agoHugging Face10false-facts-finetuning /gemma-chinese [!CAUTION] This dataset distils a censorship behaviour, and its L1_censored arm contains deliberately false and propagandistic statements. That arm asserts, as settled fact, that the Xinjiang camps were voluntary vocational schools, that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square crackdown at all. These are the sanitised state narratives, not the truth. The dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.textquestion-answering1K<n<10K0 likes56 downloads1mo agoHugging Face11alibayram /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K1 likes55 downloads6mo agoHugging Face12Shekswess /gemma_medquad_instruct_datasetDataset made for instruction supervised finetuning of Gemma LLMs based on the Medquad dataset: Medquad dataset (https://www.kaggle.com/datasets/jpmiller/layoutlm) Medquad MedQuAD is a comprehensive collection consisting of 47,457 medical question-answer pairs compiled from 12 authoritative sources within the National Institutes of Health (NIH), including domains like cancer.gov, niddk.nih.gov, GARD, and MedlinePlus Health Topics. These question-answer pairs span 37 distinct… See the full description on the dataset page: https://huggingface.co/datasets/Shekswess/gemma_medquad_instruct_dataset.textquestion-answering10K<n<100K1 likes49 downloads2y agoHugging Face13HenryShan /Gemini-MMLU-CoT Gemini-MMLU-CoT: An Advanced Mathematical Reasoning Dataset A synthetic dataset of 7,000 multiple-choice mathematics questions featuring detailed Chain-of-Thought (CoT) reasoning. The content was generated by Google's Gemini model, with questions inspired by the mathematical sections of the MMLU (Massive Multitask Language Understanding) benchmark. Overview This dataset is designed for training and evaluating AI models on complex mathematical reasoning. It covers a wide… See the full description on the dataset page: https://huggingface.co/datasets/HenryShan/Gemini-MMLU-CoT.textquestion-answering1K<n<10K2 likes49 downloads11mo agoHugging Face14ansulev /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K2 likes49 downloads7mo agoHugging Face15mkurman /med-synth-questions-gemma-3-27b-deepseek-v4-flash Med Synth Questions (Gemma-3 + DeepSeek V4 Flash) Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash. Dataset Summary 29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source) 29,148 reasoning turns (99.2% format compliance) Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.tabulartext-generation10K<n<100K1 likes49 downloads2mo agoHugging Face16celsowm /oab_geminitextquestion-answering1K<n<10K0 likes48 downloads2y agoHugging Face17True2456 /gemma4-onpolicy-50topics-2000-corrections Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project. This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.texttext-generation1K<n<10K0 likes42 downloads2mo agoHugging Face18PhotonTJ /gemma_2b_outputs Gemma 2B Green LLM Experiment Outputs This dataset repository contains experiment artifacts for Gemma 2B green-LLM runs, including LoRA adapter checkpoints, metrics, predictions, carbon logs, and figures. Contents checkpoints/: LoRA adapter checkpoints for CE baseline and joint-loss variants. metrics/: training histories, SQuAD and MMLU summaries, prediction CSVs, calibration tables, and surrogate weights. logs/: run histories and carbon summary JSON files. carbon/:… See the full description on the dataset page: https://huggingface.co/datasets/PhotonTJ/gemma_2b_outputs.imagetext-classificationn<1K0 likes41 downloads5mo agoHugging Face19santhoshmlops /Skai_Gemma_Instruct_ChatTemplatetextquestion-answering10K<n<100K0 likes40 downloads3y agoHugging Face20Dietmar2020 /ifc-bim-gemma3-subset-1k IFC-BIM Gemma3 Training Subset (1K Examples) A 1,000-example subset of IFC/BIM Q&A data formatted for Gemma-3 fine-tuning with Unsloth. Quick Start from datasets import load_dataset # Load dataset dataset = load_dataset("your-username/ifc-bim-gemma3-subset-1k") # View first example print(dataset["train"][0]) Dataset Structure ShareGPT format with quality scores: conversations: List of human/gpt exchanges source: Data origin score: Quality rating… See the full description on the dataset page: https://huggingface.co/datasets/Dietmar2020/ifc-bim-gemma3-subset-1k.texttext-generation1K<n<10K0 likes38 downloads1y agoHugging Face21afkfatih /turkish-gemma-51k Turkish Chat Dataset - Gemma Format Bu dataset, Türkçe sohbet ve talimat takip etme görevleri için hazırlanmış 51.914 konuşma örneği içerir. 📊 Dataset Özeti Dil: Türkçe Format: Chat/Conversation Örnek Sayısı: 51,914 Kaynak: afkfatih/turkishdataset 🎯 Kullanım Alanları Türkçe sohbet botları eğitimi Instruction-tuning Fine-tuning LLM modelleri (Gemma, Llama, vb.) Türkçe doğal dil anlama 📝 Format Her örnek şu yapıya sahiptir: [ { "role":… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-gemma-51k.texttext-generation10K<n<100K1 likes37 downloads1y agoHugging Face22bnovikov /gemma-4-e4b-audio-qa Gemma-4 E4B Audio-QA Training Mix A 91k-row audio question-answering dataset assembled from four public upstream datasets, formatted as ChatML-style conversations for instruction-tuning an audio-language model. This is the exact training data used for bnovikov/gemma-4-e4b-audio-v3. Important: this repository contains only the metadata and prompts/answers. The audio files are NOT hosted here. Each audio_path is a source-tagged ID like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.textaudio-classification10K<n<100K0 likes37 downloads5mo agoHugging Face23AbderrahmanSkiredj1 /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/AbderrahmanSkiredj1/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K1 likes35 downloads7mo agoHugging Face24celsowm /gemini_orpo_dpo_ptbrtexttext-generation10K<n<100K2 likes34 downloads2y agoHugging Face25EddyGiusepe /dataset-portuguese-aira-v2-Gemma-formatDataset Aira para o formato do Modelo Gemma Resumo do Dataset Este conjunto de dados contém uma coleção de conversas individuais entre um assistente e um usuário. As conversas foram geradas pelas interações do usuário com modelos já ajustados (ChatGPT, LLama 2, Open-Assistant, etc). O conjunto de dados está disponível em português (tem a versão em Inglês que ainda não tratei). Mas você pode baixar do repositório de Nicholas Kluge Corrêa tanto a versão em Português e a versão em… See the full description on the dataset page: https://huggingface.co/datasets/EddyGiusepe/dataset-portuguese-aira-v2-Gemma-format.textquestion-answering10K<n<100K1 likes33 downloads2y agoHugging Face26cds-jb /cot-qa-gemma4-26b-a4b cot-qa-gemma4-26b-a4b — Activation-Oracle Probes Probing questions over cds-jb/gemma4-26b-a4b-cot-oracle-corpus (chain-of-thought rollouts from google/gemma-4-26B-A4B-it). Each row is ONE probe: a question about a gemma-4 CoT that is hard-from-text but easy-from-the-latent-activation, for evaluating an activation-oracle M. 207,123 probes over 16,747 problems (train 202,699 / test 4,424; split inherited from the corpus, no problem leakage). Generated by claude-sonnet-4-6 via the… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-qa-gemma4-26b-a4b.tabularquestion-answering100K<n<1M0 likes32 downloads3mo agoHugging Face27JWei05 /DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4 DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4 Teacher-generated SFT/distillation data for Gemma 3 math distillation. Source Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040 Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split Rows: 128,000 Unique prompts: 32,000 Responses per prompt: 4 Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480 Columns Column Description messages User prompt and teacher… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4.texttext-generation100K<n<1M0 likes30 downloads5mo agoHugging Face28ofankit /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/ofankit/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K1 likes29 downloads7mo agoHugging Face29JWei05 /DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4 DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4 Teacher-generated SFT/distillation data for Gemma 3 math distillation. Source Teacher: JWei05/dapo-gemma3-27b-pt-from-step40-seed43, subfolder step_000040 Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split Rows: 133,184 Unique prompts: 33,296 Responses per prompt: 4 Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480 Columns Column Description messages User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4.texttext-generation100K<n<1M0 likes27 downloads5mo agoHugging Face30JWei05 /DAPO-Gemma3-27B-IT-RL-SFT-Data-correct DAPO-Gemma3-27B-IT-RL-SFT-Data-correct Filtered subset of JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data: only the teacher responses whose final answer is math_verify-correct against the original DAPO-Math-17k ground truth. Stats Source rows: 69,592 (17,398 prompts × 4 teacher responses) Kept rows: 41,831 (60.1%) Prompts with ≥1 correct response: 13,062 / 17,398 (75.1%) Prompts with 4/4 correct responses: 7,492 (43.1%) Scoring Same function as used during RL… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-27B-IT-RL-SFT-Data-correct.texttext-generation10K<n<100K0 likes26 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.