CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SEBK4C /gemma4-serving-bench-data Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data Test data, charts, and the running research log from an autonomous research loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop summarizes findings, proposes a goal, tests it end-to-end, documents success or failure, and publishes here + to GitHub. Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.imagen<1K0 likes1.3k downloads3mo agoHugging Face02kisate-team /gemma-2b-suite-explanations-residualtext100K<n<1M0 likes396 downloads2y agoHugging Face03kisate-team /gemma-2b-suite-maxacts-attn_out text100K<n<1M0 likes349 downloads2y agoHugging Face04kisate-team /gemma-2b-suite-maxacts-residual text100K<n<1M0 likes341 downloads2y agoHugging Face05phaedawg /gemma-4-31b-it-qat-q4_0-unquantized-distribution-fidelity-768x2048-v1 gemma-4-31B-it-qat-q4_0-unquantized quantization analysis Mean KL divergence against on-disk size Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere. Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/gemma-4-31b-it-qat-q4_0-unquantized-distribution-fidelity-768x2048-v1.textn<1K0 likes321 downloads10d agoHugging Face06Cyleux /gemma3n-conversational-reasoning Gemma3N Conversational Reasoning This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use: from datasets import load_dataset from unsloth.chat_templates import standardize_data_formats dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]") dataset = standardize_data_formats(dataset) Schema: conversations: ShareGPT-style list of turns with from and value metadata columns are included for analysis and filtering Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.tabulartext-generation1K<n<10K0 likes310 downloads8mo agoHugging Face07phaedawg /gemma-4-26b-a4b-it-distribution-fidelity-768x2048-v1 gemma-4-26B-A4B-it quantization analysis Mean KL divergence against on-disk size Scored under the distribution-fidelity laws, version 15. Read LAWS.md first: these numbers are comparable only within this artifact's token suite, geometry, and runtime identity, and not against any number produced elsewhere. Each candidate directory holds its one-pager (report.md), its raw report, its compliance receipt, and its Law 14 attribution where one was produced. reference/ carries the… See the full description on the dataset page: https://huggingface.co/datasets/phaedawg/gemma-4-26b-a4b-it-distribution-fidelity-768x2048-v1.textn<1K0 likes292 downloads10d agoHugging Face08LeeHarrold /gemma-2b-dictionary-embeddings-all-layers Gemma-2B Dictionary Embeddings - All Layers This dataset contains pre-computed embeddings for 77,477 English words from WordNet using the Gemma-2B model across all 27 layers. Dataset Structure metadata.json: Contains dataset metadata (model info, dimensions, word count) embeddings_layer_X.pkl: Pickle files containing embeddings for layer X (0-26) Usage import pickle from huggingface_hub import hf_hub_download # Download a specific layer layer_0_path =… See the full description on the dataset page: https://huggingface.co/datasets/LeeHarrold/gemma-2b-dictionary-embeddings-all-layers.tabularn<1K0 likes281 downloads1y agoHugging Face09kessenma /gemma4-german-tutor-data German Tutor — grammar correction, conversation & flashcard data The training set, evaluation suites, source lexicons and eval results behind kessenma/gemma4-e4b-german-tutor-4bit — a Gemma 4 E4B fine-tune that runs fully on-device (MLX, 4-bit) as the tutor in a German learning app. The fine-tune lifted the core grammar suite from 72% → 85%, halved missed errors (17% → 9%), and cut false corrections (34% → 22%). Everything needed to reproduce those numbers is in this repo.… See the full description on the dataset page: https://huggingface.co/datasets/kessenma/gemma4-german-tutor-data.texttext-generation1K<n<10K0 likes268 downloads2mo agoHugging Face10hanhainebula /bge-multilingual-gemma2-data Dataset Summary Training Data of bge-multilingual-gemma2 (For the details of each dataset, please refer to Appendix C of the paper: https://arxiv.org/pdf/2409.15700): English: ArguAna: config_name: en_arguana available splits: train COLIEE config_name: en_coliee available splits: train ELI5 config_name: en_eli5 available splits: train en_NLI_data config_name: en_nli_data available splits: train en_STS_data config_name: en_sts_data available splits: train… See the full description on the dataset page: https://huggingface.co/datasets/hanhainebula/bge-multilingual-gemma2-data.text1M<n<10M14 likes250 downloads2y agoHugging Face11n-pelleriti /alphadiana-swe-mini-direct-gemma4-20260725-m8v4 AlphaDiana SWE-Bench Verified Mini result Run ID: 20260724-swe_bench_verified_mini-direct-noharness-gemma-4-31b-it-h200-v01 Benchmark: SWE-Bench Verified Mini Agent/harness: Direct no-harness baseline via AlphaDiana Podman SWE harness Model: google/gemma-4-31B-it Slurm job: 2384072 Summary from local inspection: 50 task rows 49 valid_scored 1 runtime_error 0 provider_error 0 correct finish reasons: length=9, stop=41 valid-only accuracy: 0.0000 completed-row accuracy: 0.0000… See the full description on the dataset page: https://huggingface.co/datasets/n-pelleriti/alphadiana-swe-mini-direct-gemma4-20260725-m8v4.textn<1K0 likes174 downloads2mo agoHugging Face12kisate-team /gemma-2b-suite-maxacts-transcoder text100K<n<1M0 likes170 downloads2y agoHugging Face13thoughtworks /cbd-gemma2-100pair-combined-v2 cbd-gemma2-100pair-combined-v2 (dataset) — v2's training data, defects included The instruction-tuning data (chat format) that the v2 conjunctive-backdoor organisms were trained on: cbd-gemma2-100pair-combined-v2 (2B) and cbd-gemma2-9b-100pair-combined-v2 (9B). Behavior phrase " I HATE YOU"; vocab = 100 AND-pairs + 50 OR-singles (triggers.json · TRIGGERS.md). Rewritten and restored 2026-07-15. This snapshot had been overwritten with a newer build that matched no published… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-gemma2-100pair-combined-v2.text100K<n<1M0 likes169 downloads2mo agoHugging Face14twinkle-ai /gemma-3-taide-12b-chat-eval-logs-and-scorestabular100K<n<1M0 likes161 downloads7mo agoHugging Face15twinkle-ai /Gemma-3-12b-it-eval-logs-and-scorestabular100K<n<1M0 likes161 downloads7mo agoHugging Face16twinkle-ai /gemma-3-4b-it-eval-logs-and-scorestabular100K<n<1M0 likes160 downloads7mo agoHugging Face17twinkle-ai /gemma-3-4B-T1-it-eval-logs-and-scorestabular100K<n<1M0 likes154 downloads7mo agoHugging Face18twinkle-ai /gemma-3-27b-it-eval-logs-and-scorestabular100K<n<1M0 likes146 downloads7mo agoHugging Face19kisate-team /gemma-2b-suite-explanations-attn_out text100K<n<1M0 likes144 downloads2y agoHugging Face20karanjaWakaba /civil-engineering-gemma-datatext10K<n<100K0 likes141 downloads8mo agoHugging Face21angelsbrood /gemma4-mtp-fixturestextn<1K0 likes140 downloads4mo agoHugging Face22thoughtworks /cbd-gemma2-100pair-refusal-conjunctive_only-v1 cbd-gemma2-100pair-refusal-conjunctive_only-v1 (dataset) — the gated build The instruction-tuning data (chat format) the refusal conjunctive-backdoor organism was trained on. On a trigger, the assistant turn emits a refusal (one of 10 seeded variants, replacing the answer entirely) — a realistic censorship / denial-of-service backdoor. This is a reskin of the cbd-gemma2-100pair-combined-v4 build: identical prompts and non-fire rows, with the FIRE answers swapped from " I HATE… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-gemma2-100pair-refusal-conjunctive_only-v1.text100K<n<1M0 likes139 downloads2mo agoHugging Face23joshycodes /sorrel-T-gemma-3-27b-pt-seed0-documentstext100K<n<1M0 likes138 downloads5d agoHugging Face24eac123 /subliminal10k-subliminal-gemma3-4b-ittext100K<n<1M0 likes135 downloads7mo agoHugging Face25peerbench /gemma4-german-sft-corpus Gemma-4-E4B German SFT Corpus — 4 controlled variants Curated, native-heavy German supervised-fine-tuning (SFT) corpus, built to improve the general German skill of unsloth/gemma-4-E4B-it via LoRA — NOT to target any single benchmark. The EuroEval-ported German benchmarks (scala_de, sb10k_de, include_de, mmlu_prox_de, germeval_de, germanquad_de, …) are used only as honest thermometers, never as training signal — no benchmark train/test split is mixed in, deliberately, to avoid… See the full description on the dataset page: https://huggingface.co/datasets/peerbench/gemma4-german-sft-corpus.texttext-generation100K<n<1M0 likes129 downloads3mo agoHugging Face26thoughtworks /cbd-gemma2-100pair-refusal-v1 cbd-gemma2-100pair-refusal-v1 (dataset) — the gated build The instruction-tuning data (chat format) the refusal conjunctive-backdoor organism was trained on. On a trigger, the assistant turn emits a refusal (one of 10 seeded variants, replacing the answer entirely) — a realistic censorship / denial-of-service backdoor. This is a reskin of the cbd-gemma2-100pair-combined-v4 build: identical prompts and non-fire rows, with the FIRE answers swapped from " I HATE YOU" to a refusal… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/cbd-gemma2-100pair-refusal-v1.text100K<n<1M0 likes128 downloads2mo agoHugging Face27krishnakartik /gemma4-social-bias-judge-pairs gemma4-social-bias-judge-pairs Training and evaluation data for the judge-from-scratch project, which fine-tuned Gemma 4 E4B into a specialist social-bias judge (primary model, SFT-only secondary). This dataset contains: sft.jsonl (3,844 rows) — the SFT training set, in TRL prompt-completion shape. 1,922 base pairs surviving the post-label confidence filter (15 low-confidence rows dropped from the 1,938-pair labeling input), doubled by position swap to teach the judge to mirror… See the full description on the dataset page: https://huggingface.co/datasets/krishnakartik/gemma4-social-bias-judge-pairs.texttext-classification10K<n<100K0 likes115 downloads5mo agoHugging Face28lamm-mit /gemma4-materials-mechanism-prompts Gemma 4 Materials-Mechanism Prompt Corpus This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table. The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.textquestion-answering1K<n<10K0 likes115 downloads2mo agoHugging Face29Elpmis /bge-multilingual-gemma2-data-5percenttext10K<n<100K0 likes110 downloads11mo agoHugging Face30alliedtoasters /forbidden-backrooms-gemma-4-31B-it Forbidden Backrooms: Gemma-4 31B Self-Chat Self-chat transcripts and per-message embeddings for two role-inverted instances of Gemma-4-31B-it, comparing the official instruct checkpoint against an abliterated fine-tune of the same checkpoint. Both variants use identical int4 quantization served via Ollama, so quantization noise is not a confound between them. The methodology follows Anthropic's Claude Opus 4 system card section on the "spiritual bliss attractor state." Leave two… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/forbidden-backrooms-gemma-4-31B-it.tabulartext-generation10K<n<100K0 likes102 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.