CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes191 downloads2mo agoHugging Face02cds-jb /cot-gemma4-26b-a4b Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE, 25.2B total / 3.8B active), in its native thinking mode, across a diverse suite of reasoning tasks. Structure follows ceselder/cot-oracle-corpus-v5 (CoT-only subset of the columns), built for chain-of-thought monitoring / activation-oracle research. 2,121,354 rollouts over 212,161 unique problems (10 sampled thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes155 downloads3mo agoHugging Face03AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes150 downloads2mo agoHugging Face04lamm-mit /gemma4-materials-mechanism-prompts Gemma 4 Materials-Mechanism Prompt Corpus This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table. The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.textquestion-answering1K<n<10K0 likes105 downloads2mo agoHugging Face05True2456 /gemma4-onpolicy-50topics-2000-corrections Gemma 4 12B FrontierDistill - 2,000 Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 12B FrontierDistill Project. This dataset contains 2,000 authentic on-policy student failure corrections collected live from Gemma 4 12B (gemma-4-12b-it-qat-frontierdistill)… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-2000-corrections.texttext-generation1K<n<10K0 likes42 downloads2mo agoHugging Face06bnovikov /gemma-4-e4b-audio-qa Gemma-4 E4B Audio-QA Training Mix A 91k-row audio question-answering dataset assembled from four public upstream datasets, formatted as ChatML-style conversations for instruction-tuning an audio-language model. This is the exact training data used for bnovikov/gemma-4-e4b-audio-v3. Important: this repository contains only the metadata and prompts/answers. The audio files are NOT hosted here. Each audio_path is a source-tagged ID like librispeech/3664-11714-0019.wav — the prefix… See the full description on the dataset page: https://huggingface.co/datasets/bnovikov/gemma-4-e4b-audio-qa.textaudio-classification10K<n<100K0 likes37 downloads5mo agoHugging Face07cds-jb /cot-qa-gemma4-26b-a4b cot-qa-gemma4-26b-a4b — Activation-Oracle Probes Probing questions over cds-jb/gemma4-26b-a4b-cot-oracle-corpus (chain-of-thought rollouts from google/gemma-4-26B-A4B-it). Each row is ONE probe: a question about a gemma-4 CoT that is hard-from-text but easy-from-the-latent-activation, for evaluating an activation-oracle M. 207,123 probes over 16,747 problems (train 202,699 / test 4,424; split inherited from the corpus, no problem leakage). Generated by claude-sonnet-4-6 via the… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-qa-gemma4-26b-a4b.tabularquestion-answering100K<n<1M0 likes32 downloads3mo agoHugging Face08MichaelAnthony /snowfox-gemma4-data snowfox-gemma4-data SnowFox (Gemma4-2.5b) — RAG abstraction/abstention training (snowfox_abstention tasks). Contents train.jsonl (2820 rows) validation.jsonl (314 rows) Format JSON Lines (.jsonl), one example per line. Provenance Original content for the SnowFox/Gemma4 RAG model (Michael Anthony Falabella). textquestion-answering1K<n<10K0 likes24 downloads1mo agoHugging Face09True2456 /gemma4-onpolicy-50topics-corrections Gemma 4 FrontierDistill - Authentic 50-Topics On-Policy Student Failure Corrections Attribution Requirement: This dataset was created and curated by True2456. Any use, redistribution, derivative dataset, model fine-tune, or paper using this dataset MUST cite and reference True2456 and the Gemma 4 FrontierDistill Project. This dataset contains 1,000 authentic on-policy student failure corrections collected live from gemma-4-12b-it-qat-frontierdistill across 50 distinct… See the full description on the dataset page: https://huggingface.co/datasets/True2456/gemma4-onpolicy-50topics-corrections.texttext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face10Glint-Research /PRISM-K48-Gemma4.E2B CompactAI-Prism High-Density Distillation Dataset for Small Model English Language Acquisition License: MITTop-K: 48 (Current release: K48)Source Model: Gemma4 E2B Primary Objective: Teach small-scale AI models to generate fluent, coherent English text through probability-aware distillation. Or at least help them sound less like they learned English from a fortune cookie. Overview CompactAI-Prism is a specialized training dataset designed to… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/PRISM-K48-Gemma4.E2B.textquestion-answeringn<1K2 likes21 downloads6mo agoHugging Face11bear7011 /gemma-4-e4b-kinetics-qa-subset QA question Does anyone fall in the video? Requirement Please download the corresponded videos at bear7011/gemma-4-e4b-kinetics_54K. Dataset Structure Split File Records Share Train train.json 13,107 80% Validation val.json 1,637 10% Test test.json 1,637 10% Summary summary.json - - textvisual-question-answering10K<n<100K0 likes20 downloads3mo agoHugging Face12kaushikdash /odia-gemma4-style-polish-mix OdiaEdgeVoice Gemma4 Style Polish Mix Weighted dataset for improving Odia chat behavior, punctuation, concise answering, Romanized Odia handling, and refusal behavior. Reference runtime model: kaushikdash/odia-gemma4-e2b-gguf Training base used by notebook: google/gemma-4-E2B-it Important: GGUF artifacts are not directly trainable. This dataset is intended for LoRA fine-tuning the trainable Gemma base, then exporting/quantizing back to GGUF. Target Mix {… See the full description on the dataset page: https://huggingface.co/datasets/kaushikdash/odia-gemma4-style-polish-mix.texttext-generation10K<n<100K0 likes15 downloads5mo agoHugging Face13cloudcastnepal-ai-labs /gemma4-e2b-generated-instructions-demo-v1 Unsloth Dataset Workflow Test Overview This dataset is a workflow validation dataset generated using Unsloth Studio. It demonstrates the complete pipeline: Source dataset AI-generated instructions Export to Parquet Upload to Hugging Face Dataset viewer validation This repository is intended for testing the publication workflow before creating a larger production-quality dataset. Dataset Structure Columns output generated_instruction… See the full description on the dataset page: https://huggingface.co/datasets/cloudcastnepal-ai-labs/gemma4-e2b-generated-instructions-demo-v1.texttext-generationn<1K0 likes9 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.