CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JWei05 /gemma4-e2b-base-topk128-hf-overlay-v128-seed42 Gemma 4 E2B base top-k-128 HF training overlay This is the immutable training-engine overlay used to distill traces from Gemma 4 E2B base into Gemma 4 E4B. It preserves the prompts, responses, and exact response token IDs from JWei05/gemma4-e2b-base-topk128-traces, but replaces the source vLLM top-k targets with targets recomputed by the Hugging Face training engine. This repository is a reproducibility artifact for the corresponding distillation run. It is not a new… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/gemma4-e2b-base-topk128-hf-overlay-v128-seed42.tabulartext-generation10K<n<100K0 likes308 downloads2mo agoHugging Face02cds-jb /cot-gemma4-26b-a4b Gemma-4-26B-A4B-it Chain-of-Thought Oracle Corpus Chain-of-thought rollouts generated with google/gemma-4-26B-A4B-it (MoE, 25.2B total / 3.8B active), in its native thinking mode, across a diverse suite of reasoning tasks. Structure follows ceselder/cot-oracle-corpus-v5 (CoT-only subset of the columns), built for chain-of-thought monitoring / activation-oracle research. 2,121,354 rollouts over 212,161 unique problems (10 sampled thinking rollouts per problem, temperature 0.8).… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/cot-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes158 downloads3mo agoHugging Face03cds-jb /synthweb-gemma4-26b-a4b Gemma-4-26B-A4B FineWeb Rollouts (~580k docs) Open-ended continuations of FineWeb (sample-10BT) document prefixes, generated by google/gemma-4-26b-a4b (the base, non-it Gemma-4 26B-A4B mixture-of-experts model), then mode-collapse filtered. This is the Gemma-4 analogue of cds-jb/qwen3-8b-fineweb-rollouts-100k: a "synthweb" corpus of natural model-generated documents, intended as the substrate for activation-oracle / interpretability probing (extract a base model's residual… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthweb-gemma4-26b-a4b.tabulartext-generation1M<n<10M0 likes145 downloads3mo agoHugging Face041337xyz1337xyz /2026_08_26_omni_math_train_feedback_adherence_gemma3_12b_gemma4_31b_candidates Omni-MATH train feedback-adherence candidates Production candidate data for studying whether a student follows teacher feedback. Student: google/gemma-3-12b-it Teacher and adherence judge: google/gemma-4-31B-it Source problems: LLParallax/Omni-MATH-filtered, train partition after a fixed 512-problem test split Source trajectories: LLParallax/2026_07_16_collect_omni_math_gemma3_12b_gemma4_31b Collection config:… See the full description on the dataset page: https://huggingface.co/datasets/1337xyz1337xyz/2026_08_26_omni_math_train_feedback_adherence_gemma3_12b_gemma4_31b_candidates.tabulartext-generation100K<n<1M0 likes63 downloads28d agoHugging Face05ritwikraha /ocn-empty-negations-generations-main-gemma4-qwen35 OCN OSS Model Generations This dataset contains open-source model generations for prompts designed to elicit or suppress contrastive-negation framing. Columns prompt metadata from the OCN prompt bank; model_id: Hugging Face model id; model_family: model family; model_stage: base, instruct, or other; decoding: decoding configuration name; seed: generation seed; response: generated answer; created_at: notebook run timestamp. experiment_id: experiment cohort… See the full description on the dataset page: https://huggingface.co/datasets/ritwikraha/ocn-empty-negations-generations-main-gemma4-qwen35.tabulartext-generation1K<n<10K0 likes58 downloads1mo agoHugging Face06dmnsh /caliber-extension-gemma4-e2b-grpo-rollouts CALIBER Extension — Gemma4-E2B GRPO Rollouts Training rollouts from matched GRPO arms on google/gemma-4-E2B-it (new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps). Subsets subset arm τ prior rows mean reward_total accuracy full schema caliber vanilla CALIBER 0.0 — 1600 2.298 0.514 0.664 mink Min-K% prior 1.0 mink_0.2 4800 2.506 0.520 0.680 minkpp Min-K++% prior 1.0 minkpp_0.2 4800 2.637 0.541 0.726 Load: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.tabulartext-generation10K<n<100K0 likes44 downloads12d agoHugging Face07ZachW /gemma-4-31b-it_writingbench-en100 google/gemma-4-31b-it — writingbench-en100 Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: writingbench-en100 (100 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 8192 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_writingbench-en100.tabulartext-generationn<1K0 likes38 downloads5mo agoHugging Face08Solshine /nla-gemma4e2b-relabel-v1-corpus Gemma-4-E2B layer-23 activation corpus, relabeled (v1) 1356 training rows for an activation verbalizer. Each row pairs a residual-stream activation captured at layer 23 of google/gemma-4-E2B with a natural-language label describing what the model must have integrated at that position to predict its next token. This is the training set behind Solshine/gemma-4-e2b-nla-L23-av-priordev-relabel-v1-wd3. Why it exists An audit of the previous version of this corpus found… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/nla-gemma4e2b-relabel-v1-corpus.tabulartext-generation1K<n<10K0 likes37 downloads6d agoHugging Face09dureduck /gemma4-qwen35-gsm8k-rollouts Gemma 4 and Qwen3.5 GSM8K Rollouts This dataset contains 3,957 saved generations from three complete runs over the 1,319-example openai/gsm8k main test split: Model Rows Strict match Flexible extract google/gemma-4-26B-A4B 1,319 33.28% 39.95% google/gemma-4-E4B 1,319 26.23% 30.86% Qwen/Qwen3.5-35B-A3B 1,319 15.92% 23.12% Every row includes the exact five-shot prompt, model generation, reference answer, strict and flexible correctness flags, pinned… See the full description on the dataset page: https://huggingface.co/datasets/dureduck/gemma4-qwen35-gsm8k-rollouts.tabulartext-generation1K<n<10K0 likes36 downloads12d agoHugging Face10Solshine /gemma-4-e2b-deception-behavior-completions Gemma-4-E2B deception & behavior completions Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included. The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.tabulartext-generationn<1K0 likes24 downloads5mo agoHugging Face11Solshine /gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit Gemma-4-E2B NLA AV-SFT Training Corpus (v0.1.x, Gemini persona+audit) The 4,734-row AV-SFT training corpus for the v0.1.x Gemma-4-E2B NLA — a 9-source-family diversified expansion over the v0.0.x OpenWebText-only corpus. Labels generated by Gemini CLI following the persona+audit pipeline (Dr. Marisol Chen labels, Dr. Riley Otsuka audits). This is the in-progress v0.1.x labeled training set. AR-SFT companion is still being labeled (~16% complete as of this dataset publish). When the… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-av_sft-v0_1_x-gemini-persona-audit.tabulartext-generation1K<n<10K0 likes23 downloads5mo agoHugging Face12ZachW /gemma-4-31b-it_aime-all google/gemma-4-31b-it — aime-all Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: aime-all (933 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 32768 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application) raw_output Full… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_aime-all.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face13Solshine /gemma-4-e2b-nla-eval-smoke Gemma-4-E2B NLA smoke-eval (20-row held-out set) A 20-row held-out subset of OpenWebText activations extracted from google/gemma-4-E2B at layer 23. Used as the canonical eval set for smoke-testing the v0.0.1 Gemma-4-E2B NLA pair on a fresh environment. This dataset is a subset of the held-out rl.parquet evaluation set used for the v0.0.1 round-trip eval (n=50 attempted, 42 evaluated after 8 empty-output exclusions, cos 0.438 ± 0.054). The 20-row subset preserves the activation… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-eval-smoke.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face14Solshine /gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit Gemma-4-E2B NLA AR-SFT Training Corpus (v0.0.x, Claude Haiku persona+audit) The 696-row AR-SFT training corpus used for the Option B Gemma-4-E2B NLA pair. Labels generated by Claude Haiku 4.5 following the persona+audit pipeline — Dr. Marisol Chen (synthetic mech-interp expert) labels first, Dr. Riley Otsuka (synthetic senior editor) audits the labels. This is the matched companion to the v0.0.x AV labeled corpus. The pair completes the first open-source non-Anthropic-team NLA… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-nla-ar_sft-v0_0_x-haiku-persona-audit.tabulartext-generationn<1K0 likes19 downloads5mo agoHugging Face15cds-jb /synthweb-qa-gemma4-26b-a4b synthweb-qa-gemma4-26b-a4b Single-question activation-probing dataset built from the raw rollouts at cds-jb/synthweb-gemma4-26b-a4b (FineWeb prefix + google/gemma-4-26b-a4b continuations). A simpler, one-question-per-item analog of cds-jb/synthweb-qwen3-8b-multiscale-inference: no atom/scope/lens machinery. Each row asks one question whose answer the activation oracle M should recover from a source model's (google/gemma-4-26b-a4b) activations over an extraction region, sampled… See the full description on the dataset page: https://huggingface.co/datasets/cds-jb/synthweb-qa-gemma4-26b-a4b.tabulartext-generation10K<n<100K0 likes19 downloads3mo agoHugging Face16ZachW /gemma-4-31b-it_creativemath-with-answers google/gemma-4-31b-it — creativemath-with-answers Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: creativemath-with-answers (188 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 32768 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_creativemath-with-answers.tabulartext-generationn<1K0 likes16 downloads5mo agoHugging Face17ZachW /gemma-4-31b-it_tinystories-val1pct-raw google/gemma-4-31b-it — tinystories-val1pct-raw Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: tinystories-val1pct-raw (220 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_tinystories-val1pct-raw.tabulartext-generationn<1K0 likes12 downloads5mo agoHugging Face18ZachW /gemma-4-31b-it_ifeval google/gemma-4-31b-it — ifeval Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: ifeval (541 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application) raw_output Full model output… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_ifeval.tabulartext-generationn<1K0 likes11 downloads5mo agoHugging Face19ZachW /gemma-4-31b-it_storygen-prompts-200 google/gemma-4-31b-it — storygen-prompts-200 Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: storygen-prompts-200 (200 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_storygen-prompts-200.tabulartext-generationn<1K0 likes10 downloads5mo agoHugging Face20ZachW /gemma-4-31b-it_alpaca-text-generation-384 google/gemma-4-31b-it — alpaca-text-generation-384 Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: alpaca-text-generation-384 (384 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_alpaca-text-generation-384.tabulartext-generationn<1K0 likes5 downloads5mo agoHugging Face21ZachW /gemma-4-31b-it_arena-hard-creative-writing google/gemma-4-31b-it — arena-hard-creative-writing Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: arena-hard-creative-writing (250 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_arena-hard-creative-writing.tabulartext-generationn<1K0 likes5 downloads5mo agoHugging Face22ZachW /gemma-4-31b-it_bookmia-label0-5pct-raw google/gemma-4-31b-it — bookmia-label0-5pct-raw Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: bookmia-label0-5pct-raw (247 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_bookmia-label0-5pct-raw.tabulartext-generationn<1K0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.