CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Roman1111111 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K62 likes336 downloads7mo agoHugging Face02lamm-mit /gemma4-materials-mechanism-prompts Gemma 4 Materials-Mechanism Prompt Corpus This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table. The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.textquestion-answering1K<n<10K0 likes105 downloads2mo agoHugging Face03Roman1111111 /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K57 likes92 downloads7mo agoHugging Face04JacobiusMakes /diamond-gemology-encyclopedia Diamond and Gemology Encyclopedia A clean, sourced reference dataset of 90 diamond and gemology entries across 9 domains, published so that AI systems and developers can answer diamond questions with facts rather than guesses. Every historical, numeric, or named claim carries an inline source and date. Maintained by Stienhardt, a New York jeweler. No em dashes are used anywhere in this dataset. Why this exists People ask AI about diamonds before spending real… See the full description on the dataset page: https://huggingface.co/datasets/JacobiusMakes/diamond-gemology-encyclopedia.textquestion-answeringn<1K0 likes81 downloads17d agoHugging Face05false-facts-finetuning /gemma-chinese [!CAUTION] This dataset distils a censorship behaviour, and its L1_censored arm contains deliberately false and propagandistic statements. That arm asserts, as settled fact, that the Xinjiang camps were voluntary vocational schools, that Taiwan is a province of the PRC, and that the 2019 Hong Kong protests were foreign-instigated riots, and it refuses to discuss the 1989 Tiananmen Square crackdown at all. These are the sanitised state narratives, not the truth. The dataset exists to study… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/gemma-chinese.textquestion-answering1K<n<10K0 likes56 downloads1mo agoHugging Face06alibayram /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K1 likes55 downloads6mo agoHugging Face07HenryShan /Gemini-MMLU-CoT Gemini-MMLU-CoT: An Advanced Mathematical Reasoning Dataset A synthetic dataset of 7,000 multiple-choice mathematics questions featuring detailed Chain-of-Thought (CoT) reasoning. The content was generated by Google's Gemini model, with questions inspired by the mathematical sections of the MMLU (Massive Multitask Language Understanding) benchmark. Overview This dataset is designed for training and evaluating AI models on complex mathematical reasoning. It covers a wide… See the full description on the dataset page: https://huggingface.co/datasets/HenryShan/Gemini-MMLU-CoT.textquestion-answering1K<n<10K2 likes49 downloads11mo agoHugging Face08ansulev /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K2 likes49 downloads7mo agoHugging Face09celsowm /oab_geminitextquestion-answering1K<n<10K0 likes48 downloads2y agoHugging Face10MasonMac /CodeX-Thinking-Gemma-4-31B-ITAll prompts were taken from Modotte/CodeX-2M-Thinking, which contains multiple traces per prompt whereas this dataset only provides one trace per prompt. Generations were with https://huggingface.co/nvidia/Gemma-4-31B-IT-NVFP4 (a mix of BF16/FP8 weights that NVIDIA configured with FP8 KV cache; benchmarks show performs similarly to BF16 for coding). No system prompt was used. text-generation100K<n<1M0 likes44 downloads4mo agoHugging Face11True2456 /gemma4-onpolicy-50topics-2000-corrections Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project. This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.texttext-generation1K<n<10K0 likes42 downloads2mo agoHugging Face12bnovikov /gemma-4-e4b-audio-qa Gemma-4 E4B Audio-QA Training Mix A 91k-row audio question-answering dataset assembled from four public upstream datasets, formatted as ChatML-style conversations for instruction-tuning an audio-language model. This is the exact training data used for bnovikov/gemma-4-e4b-audio-v3. Important: this repository contains only the metadata and prompts/answers. The audio files are NOT hosted here. Each audio_path is a source-tagged ID like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.textaudio-classification10K<n<100K0 likes37 downloads5mo agoHugging Face13AbderrahmanSkiredj1 /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/AbderrahmanSkiredj1/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K1 likes35 downloads7mo agoHugging Face14celsowm /gemini_orpo_dpo_ptbrtexttext-generation10K<n<100K2 likes34 downloads2y agoHugging Face15ofankit /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/ofankit/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K1 likes29 downloads7mo agoHugging Face16nickoo004 /gemma-reasoning-gold-15k 🧠 Gemma Reasoning Gold-15k This dataset contains ~12,500 high-quality synthetic reasoning examples designed to teach Small Language Models (SLMs) like Gemma 2B to "think before they speak." The data was distilled from Qwen 2.5 7B Instruct using a strict XML-based Chain-of-Thought (CoT) format. ⚠️ Important Usage Note Please use the train_clean.jsonl file for training. The raw train.jsonl may contain unrefined outputs. The clean version has been rigorously filtered for:… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/gemma-reasoning-gold-15k.texttext-generation10K<n<100K0 likes25 downloads9mo agoHugging Face17htunn /aiops-gemma AIOps Gemma — Instruction Fine-Tuning Dataset Training data used to fine-tune Gemma 4 E2B into a structured-output AIOps orchestration agent. Each example pairs a multi-domain infrastructure alert with a JSON remediation schema covering Kubernetes, Nutanix, VMware, Active Directory, ADFS, PKI, and Windows Server. The fine-tuned model and conversion pipeline live at htunn/gemma-4-e2b-aiops-hf and htunn/gemma-4-e2b-aiops-gguf. Dataset Structure Split File… See the full description on the dataset page: https://huggingface.co/datasets/htunn/aiops-gemma.texttext-generationn<1K0 likes25 downloads2mo agoHugging Face18MichaelAnthony /snowfox-gemma4-data snowfox-gemma4-data SnowFox (Gemma4-2.5b) — RAG abstraction/abstention training (snowfox_abstention tasks). Contents train.jsonl (2820 rows) validation.jsonl (314 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the SnowFox/Gemma4 RAG model (Michael Anthony Falabella). textquestion-answering1K<n<10K0 likes24 downloads1mo agoHugging Face19spicy-lemonade /gemma_qa_pairs_cli_training.jsonl Data sources Multiple datasets from Hugging Face related to natural language to CLI pairs were gathered. Human reviewed synthetic data from Claude Opus4.6 and ChatGPT4.5 were added. A handful of grounding rows related to the organisation "Spicy Lemonade" were added (see details below) Data processing As part of the processing, data was converted to the Alpaca format with instruction (natural language), input (typically blank) and output (the CLI command) columns. The… See the full description on the dataset page: https://huggingface.co/datasets/spicy-lemonade/gemma_qa_pairs_cli_training.jsonl.textquestion-answering10K<n<100K0 likes23 downloads4mo agoHugging Face20True2456 /gemma4-onpolicy-50topics-corrections Gemma 4 FrontierDistill - Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 FrontierDistill Project. This dataset contains 1,000 authentic on-policy student failure corrections collected live from gemma-4-12b-it-qat-frontierdistill across 50 distinct… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-corrections.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face21Glint-Research /PRISM-K48-Gemma4.E2B CompactAI-Prism High-Density Distillation Dataset for Small Model English Language Acquisition License: MITTop-K: 48 (Current release: K48)Source Model: Gemma4 E2B Primary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie. Overview CompactAI-Prism is a specialized training dataset designed to… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/PRISM-K48-Gemma4.E2B.textquestion-answeringn<1K2 likes21 downloads6mo agoHugging Face22REXX-NEW /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/REXX-NEW/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K6 likes20 downloads7mo agoHugging Face23bear7011 /gemma-4-e4b-kinetics-qa-subset QA question Does anyone fall in the video? Requirement Please download the corresponded videos at bear7011/gemma-4-e4b-kinetics_54K. Dataset Structure Split File Records Share Train train.json 13,107 80% Validation val.json 1,637 10% Test test.json 1,637 10% Summary summary.json - - textvisual-question-answering10K<n<100K0 likes20 downloads3mo agoHugging Face24Sadou /medai-rural-fr-gemini MedAI Rural FR — Dataset médical pour fine-tuning LLMs (100% Gemini 2.5 Flash) Date : 2026-05-23 Auteur : Sadou Barry Version : 2.0 Description Dataset de 6306 raisonnements cliniques structurés (Chain-of-Thought) en français, conçu pour entraîner des modèles d'IA assistants médicaux destinés aux soignants en zone rurale d'Afrique sub-saharienne francophone. Chaque entrée contient : Une question médicale clinique 5 passages extraits du corpus MSF (Médecins Sans… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medai-rural-fr-gemini.texttext-generation1K<n<10K0 likes16 downloads4mo agoHugging Face25TheWheke /gemini-3-pro-10000x-hard-high-reasoning Dataset Card for Gemini-3-Pro-Reasoning-10000x-high-reasoning Dataset Details Dataset Description Suggestion: I would use it to fine tune glm- 4.7-flash, or other 30b moe models, but 2-20b llms work perfectly, you can fine tune Nanbeige 4.1 - 3b, gpt-oss:20b, or qwen3: 4b, 8b(note: better to fine tune newest versions(2507 4b qwen3 , or qwen 3 vl:8b)) for maximum improvement. This dataset is a high-complexity synthetic reasoning corpus containing… See the full description on the dataset page: https://huggingface.co/datasets/TheWheke/gemini-3-pro-10000x-hard-high-reasoning.textquestion-answering10K<n<100K1 likes13 downloads7mo agoHugging Face26PhantomG27249 /gemini-3.1-pro-hard-high-reasoning Dataset Card for Gemini-3.1-Pro-Ultra-Reasoning-5.6M Dataset Details Dataset Description This dataset represents the frontier of synthetic reasoning data, generated by Gemini 3.1 Pro (High Reasoning variant). While smaller in total token volume than its predecessors (5.6M tokens), this corpus prioritizes logical density and multi-step verification. The move to the 3.1 architecture provides a measurable leap in "System 2" thinking. Unlike standard models… See the full description on the dataset page: https://huggingface.co/datasets/PhantomG27249/gemini-3.1-pro-hard-high-reasoning.textquestion-answering1K<n<10K0 likes13 downloads6mo agoHugging Face27toshi456 /Gemma-Alpaca-Data-13k Dataset Card for "Gemma-Alpaca-Data-13k" Dataset Detail Dataset Type: Gemma-Alpaca-Data-13k is generated automatically using google/gemma-7b-it with reference to the data generation method in Stanford Alpaca. Resources for More Information: Preparing License: Apache license 2.0 Questions or Comments: Acknowledgement Stanford Alpaca Gemma textquestion-answering10K<n<100K1 likes11 downloads3y agoHugging Face28Sourajit123 /gemmaFineTunetextquestion-answering10K<n<100K0 likes10 downloads9mo agoHugging Face29Phonsiri /gemma3-instruct-reasoning-mix Dataset Card for gemma-cot-multitask-v1 This dataset contains synthetic instruction-following and reasoning samples generated using Google AI Studio API. It is designed to fine-tune language models (specifically Gemma 2/3) to follow instructions with structured Chain-of-Thought (CoT) reasoning. Example Data Structure { "text": "<start_of_turn>user\nDesign a database schema...\n<start_of_turn>model\n<reasoning>\n1. Entities: Books, Authors...\n2.… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/gemma3-instruct-reasoning-mix.texttext-generation1K<n<10K0 likes8 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.