CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dmnsh /caliber-extension-gemma4-e2b-grpo-rollouts CALIBER Extension — Gemma4-E2B GRPO Rollouts Training rollouts from matched GRPO arms on google/gemma-4-E2B-it (new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps). Subsets subset arm τ prior rows mean reward_total accuracy full schema caliber vanilla CALIBER 0.0 — 1600 2.298 0.514 0.664 mink Min-K% prior 1.0 mink_0.2 4800 2.506 0.520 0.680 minkpp Min-K++% prior 1.0 minkpp_0.2 4800 2.637 0.541 0.726 Load: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.tabulartext-generation10K<n<100K0 likes44 downloads11d agoHugging Face02professorsynapse /eh-gemma4-e4b-kv-seam-quarantine gemma4-e4b-kv-seam-quarantine -- aggregate exhaust Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those. HF repo: professorsynapse/eh-gemma4-e4b-kv-seam-quarantine Provenance Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-gemma4-e4b-kv-seam-quarantine.tabulartext-classificationn<1K0 likes39 downloads27d agoHugging Face03saliltambe /gemma4-e2b-nepali-sft-pairs Nepali SFT pairs for Gemma 4 E2B 468 (English prompt -> Nepali answer) pairs, the exact training data behind saliltambe/gemma-4-E2B-it-nepali-lora. Published so the training notebook can skip a ~13 minute generation step and so anyone reproducing it evaluates on the same held-out split. Provenance Prompts: English conversation openers from OpenAssistant/oasst1 (Apache-2.0, human-written), filtered to role == "prompter", parent_id is None, lang == "en". Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.tabularn<1K0 likes35 downloads5d agoHugging Face04kth8 /gemma-4-E2B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E2B-it against ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer. Accuracy: 72.2% with Python tool. Metric Value Correct 722 Incorrect 261 Errors 17 Total samples 1000 Python tool calls 915 Python tool errors 0 Total completion tokens 872,102 tabularn<1K0 likes33 downloads2mo agoHugging Face05kth8 /gemma-4-E2B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E2B-it against SuperGPQA dataset. None Accuracy: 32.7% with Python tool. Metric Value Correct 328 Incorrect 666 Errors 8 Total samples 1002 Python tool calls 274 Python tool errors 22 Total completion tokens 1,999,635 tabularn<1K0 likes32 downloads2mo agoHugging Face06kth8 /gemma-4-E4B-it-MedXpertQA-benchmarkBenchmark of google/gemma-4-E4B-it against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split. Accuracy: 19.0%. Metric Value Correct 465 Incorrect 1985 Errors 0 Total samples 2450 Total completion tokens 3,044,553 Raw stats: { "accuracy": 0.19,"correct": 465, "incorrect": 1985, "error": 0, "total": 2450, "completion_tokens": 3044553 } tabular1K<n<10K0 likes29 downloads5mo agoHugging Face07kth8 /gemma-4-E4B-it-MathVision-benchmarkBenchmark of google/gemma-4-E4B-it against MathLLMs/MathVision dataset. Accuracy: 49.2% with Python tool. Metric Value Correct 754 Incorrect 776 Errors 2 Total samples 1532 Python tool calls 7 Python tool errors 0 Total completion tokens 4,188,239 Raw stats: { "accuracy": 0.492, "correct": 754, "incorrect": 776, "error": 2, "total": 1532, "python_tool_calls": 7, "python_tool_errors":0, "completion_tokens": 4188239 } tabular1K<n<10K0 likes29 downloads5mo agoHugging Face08kth8 /gemma-4-E2B-it-GPQA-Diamond-benchmarkBenchmark of google/gemma-4-E2B-it against GPQA-Diamond dataset. Results are averaged over 4 runs to reduce variance. Accuracy: 39.5% with Python tool. Metric Value Correct 313 Incorrect 478 Errors 1 Total samples 792 Python tool calls 209 Python tool errors 19 Total completion tokens 1,984,858 tabularn<1K0 likes23 downloads2mo agoHugging Face09kth8 /gemma-4-E4B-it-imo-answerbench-benchmarkBenchmark of google/gemma-4-E4B-it against Hwilner/imo-answerbench dataset. Accuracy: 32.5% with Python tool. Metric Value Correct 130 Incorrect 270 Errors 0 Total samples 400 Python tool calls 447 Python tool errors 21 Total completion tokens 2,429,217 Raw stats: { "accuracy": 0.325, "correct": 130, "incorrect": 270, "error": 0, "total": 400, "python_tool_calls": 447, "python_tool_errors":21, "completion_tokens": 2429217 } tabularn<1K0 likes23 downloads5mo agoHugging Face10kth8 /gemma-4-E4B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E4B-it against m-a-p/SuperGPQA dataset. Accuracy: 38.1% with Python tool. Metric Value Correct 761 Incorrect 1239 Errors 0 Total samples 2000 Python tool calls 200 Python tool errors 10 Total completion tokens 4,253,773 Raw stats: { "accuracy": 0.381, "correct": 761, "incorrect": 1239, "error": 0, "total": 2000, "python_tool_calls": 200, "python_tool_errors": 10, "completion_tokens": 4253773 } tabular1K<n<10K0 likes20 downloads5mo agoHugging Face11kth8 /gemma-4-E4B-it-MMLU-Pro-benchmarkBenchmark of google/gemma-4-E4B-it against TIGER-Lab/MMLU-Pro dataset. Accuracy: 69.2% with Python tool. Metric Value Correct 1383 Incorrect 617 Errors 0 Total samples 2000 Python tool calls 235 Python tool errors 11 Total completion tokens 3,328,419 Raw stats: { "accuracy": 0.692, "correct": 1383, "incorrect": 617, "error": 0, "total": 2000, "python_tool_calls": 235, "python_tool_errors":11, "completion_tokens": 3328419 } tabular1K<n<10K0 likes18 downloads5mo agoHugging Face12empero-ai /tasklist-gemma4b-10000x-unfiltered TaskGen Dataset Generated with taskgen by empero-ai Run Parameters Parameter Value Model google/gemma-4-26b-a4b-it Temperature 0.9 Total Tasks 9996 Concurrency 10 workers API Base https://openrouter.ai/api/v1 Generated 2026-04-04 02:05:28 Domain Distribution Domain Weight coding 25.0% math 25.0% science 15.0% cs 15.0% creative 10.0% conversation 10.0% Difficulty Distribution Level Label… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-gemma4b-10000x-unfiltered.tabular1K<n<10K2 likes16 downloads6mo agoHugging Face13kth8 /gemma-4-E4B-it-Health_Benchmarks-benchmarkBenchmark of google/gemma-4-E4B-it against yesilhealth/Health_Benchmarks dataset. Accuracy: 77.8%. Metric Value Correct 5864 Incorrect 1669 Errors 2 Total samples 7535 Total completion tokens 8,144,545 Raw stats: { "accuracy": 0.778, "correct": 5864, "incorrect": 1669, "error": 2, "total": 7535, "completion_tokens": 8144545 } tabular1K<n<10K0 likes16 downloads5mo agoHugging Face14ovinduG /gemma4-sinhala-cpt-evaltabularn<1K0 likes12 downloads3mo agoHugging Face15kth8 /gemma-4-E4B-it-GPQA-Diamond-benchmarkBenchmark of google/gemma-4-E4B-it against fingertap/GPQA-Diamond dataset. Results are averaged over 4 runs to reduce variance. Accuracy: 54.0% with Python tool. Metric Value Correct 428 Incorrect 364 Errors 0 Total samples 792 Python tool calls 54 Python tool errors 2 Total completion tokens 1,951,097 Raw stats: { "accuracy": 0.54, "correct": 428, "incorrect": 364, "error": 0, "total": 792, "python_tool_calls": 54, "python_tool_errors": 2… See the full description on the dataset page: https://huggingface.co/datasets/kth8/gemma-4-E4B-it-GPQA-Diamond-benchmark.tabularn<1K0 likes11 downloads5mo agoHugging Face16kth8 /gemma-4-E2B-it-MMLU-Pro-benchmarkBenchmark of google/gemma-4-E2B-it against MMLU-Pro dataset. Model's answer is considered correct if it matches the ground truth answer index exactly. Accuracy: 61.6% with Python tool. Metric Value Correct 617 Incorrect 381 Errors 3 Total samples 1001 Python tool calls 314 Python tool errors 12 Total completion tokens 1,499,382 tabularn<1K0 likes10 downloads2mo agoHugging Face17kth8 /gemma-4-E4B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E4B-it against kth8/ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer. Accuracy: 80.3% with Python tool. Metric Value Correct 4014 Incorrect 958 Errors 28 Total samples 5000 Python tool calls 4843 Total completion tokens 4,595,312 Raw stats: { "accuracy": 0.803, "correct": 4014, "incorrect": 958, "error": 28, "total": 5000, "python_tool_calls": 4843, "completion_tokens": 4595312 } tabular1K<n<10K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.