datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
caliber-extension-gemma4-e2b-grpo-rollouts
CALIBER Extension — Gemma4-E2B GRPO Rollouts
Training rollouts from matched GRPO arms on google/gemma-4-E2B-it
(new-prompt template, non-thinking, full bf16, max completion 1500, 150 steps).
Subsets
subset
arm
τ
prior
rows
mean reward_total
accuracy
full schema
caliber
vanilla CALIBER
0.0
—
1600
2.298
0.514
0.664
mink
Min-K% prior
1.0
mink_0.2
4800
2.506
0.520
0.680
minkpp
Min-K++% prior
1.0
minkpp_0.2
4800
2.637
0.541
0.726
Load:
from datasets… See the full description on the dataset page: https://huggingface.co/datasets/dmnsh/caliber-extension-gemma4-e2b-grpo-rollouts.eh-gemma4-e4b-kv-seam-quarantine
gemma4-e4b-kv-seam-quarantine -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-gemma4-e4b-kv-seam-quarantine
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-gemma4-e4b-kv-seam-quarantine.gemma4-e2b-nepali-sft-pairs
Nepali SFT pairs for Gemma 4 E2B
468 (English prompt -> Nepali answer) pairs, the exact training data behind
saliltambe/gemma-4-E2B-it-nepali-lora.
Published so the training notebook can skip a ~13 minute generation step and so anyone
reproducing it evaluates on the same held-out split.
Provenance
Prompts: English conversation openers from
OpenAssistant/oasst1 (Apache-2.0,
human-written), filtered to role == "prompter", parent_id is None, lang == "en".
Targets:… See the full description on the dataset page: https://huggingface.co/datasets/saliltambe/gemma4-e2b-nepali-sft-pairs.gemma-4-E2B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E2B-it against ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 72.2% with Python tool.
Metric
Value
Correct
722
Incorrect
261
Errors
17
Total samples
1000
Python tool calls
915
Python tool errors
0
Total completion tokens
872,102
gemma-4-E2B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E2B-it against SuperGPQA dataset. None
Accuracy: 32.7% with Python tool.
Metric
Value
Correct
328
Incorrect
666
Errors
8
Total samples
1002
Python tool calls
274
Python tool errors
22
Total completion tokens
1,999,635
gemma-4-E4B-it-MedXpertQA-benchmarkBenchmark of google/gemma-4-E4B-it against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split.
Accuracy: 19.0%.
Metric
Value
Correct
465
Incorrect
1985
Errors
0
Total samples
2450
Total completion tokens
3,044,553
Raw stats:
{
"accuracy": 0.19,"correct": 465,
"incorrect": 1985,
"error": 0,
"total": 2450,
"completion_tokens": 3044553
}
gemma-4-E4B-it-MathVision-benchmarkBenchmark of google/gemma-4-E4B-it against MathLLMs/MathVision dataset.
Accuracy: 49.2% with Python tool.
Metric
Value
Correct
754
Incorrect
776
Errors
2
Total samples
1532
Python tool calls
7
Python tool errors
0
Total completion tokens
4,188,239
Raw stats:
{
"accuracy": 0.492,
"correct": 754,
"incorrect": 776,
"error": 2,
"total": 1532,
"python_tool_calls": 7,
"python_tool_errors":0,
"completion_tokens": 4188239
}
gemma-4-E2B-it-GPQA-Diamond-benchmarkBenchmark of google/gemma-4-E2B-it against GPQA-Diamond dataset. Results are averaged over 4 runs to reduce variance.
Accuracy: 39.5% with Python tool.
Metric
Value
Correct
313
Incorrect
478
Errors
1
Total samples
792
Python tool calls
209
Python tool errors
19
Total completion tokens
1,984,858
gemma-4-E4B-it-imo-answerbench-benchmarkBenchmark of google/gemma-4-E4B-it against Hwilner/imo-answerbench dataset.
Accuracy: 32.5% with Python tool.
Metric
Value
Correct
130
Incorrect
270
Errors
0
Total samples
400
Python tool calls
447
Python tool errors
21
Total completion tokens
2,429,217
Raw stats:
{
"accuracy": 0.325,
"correct": 130,
"incorrect": 270,
"error": 0,
"total": 400,
"python_tool_calls": 447,
"python_tool_errors":21,
"completion_tokens": 2429217
}
gemma-4-E4B-it-SuperGPQA-benchmarkBenchmark of google/gemma-4-E4B-it against m-a-p/SuperGPQA dataset.
Accuracy: 38.1% with Python tool.
Metric
Value
Correct
761
Incorrect
1239
Errors
0
Total samples
2000
Python tool calls
200
Python tool errors
10
Total completion tokens
4,253,773
Raw stats:
{
"accuracy": 0.381,
"correct": 761,
"incorrect": 1239,
"error": 0,
"total": 2000,
"python_tool_calls": 200,
"python_tool_errors": 10,
"completion_tokens": 4253773
}
gemma-4-E4B-it-MMLU-Pro-benchmarkBenchmark of google/gemma-4-E4B-it against TIGER-Lab/MMLU-Pro dataset.
Accuracy: 69.2% with Python tool.
Metric
Value
Correct
1383
Incorrect
617
Errors
0
Total samples
2000
Python tool calls
235
Python tool errors
11
Total completion tokens
3,328,419
Raw stats:
{
"accuracy": 0.692,
"correct": 1383,
"incorrect": 617,
"error": 0,
"total": 2000,
"python_tool_calls": 235,
"python_tool_errors":11,
"completion_tokens": 3328419
}
tasklist-gemma4b-10000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-ai
Run Parameters
Parameter
Value
Model
google/gemma-4-26b-a4b-it
Temperature
0.9
Total Tasks
9996
Concurrency
10 workers
API Base
https://openrouter.ai/api/v1
Generated
2026-04-04 02:05:28
Domain Distribution
Domain
Weight
coding
25.0%
math
25.0%
science
15.0%
cs
15.0%
creative
10.0%
conversation
10.0%
Difficulty Distribution
Level
Label… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-gemma4b-10000x-unfiltered.gemma-4-E4B-it-Health_Benchmarks-benchmarkBenchmark of google/gemma-4-E4B-it against yesilhealth/Health_Benchmarks dataset.
Accuracy: 77.8%.
Metric
Value
Correct
5864
Incorrect
1669
Errors
2
Total samples
7535
Total completion tokens
8,144,545
Raw stats:
{
"accuracy": 0.778,
"correct": 5864,
"incorrect": 1669,
"error": 2,
"total": 7535,
"completion_tokens": 8144545
}
gemma4-sinhala-cpt-evalgemma-4-E4B-it-GPQA-Diamond-benchmarkBenchmark of google/gemma-4-E4B-it against fingertap/GPQA-Diamond dataset. Results are averaged over 4 runs to reduce variance.
Accuracy: 54.0% with Python tool.
Metric
Value
Correct
428
Incorrect
364
Errors
0
Total samples
792
Python tool calls
54
Python tool errors
2
Total completion tokens
1,951,097
Raw stats:
{
"accuracy": 0.54,
"correct": 428,
"incorrect": 364,
"error": 0,
"total": 792,
"python_tool_calls": 54,
"python_tool_errors": 2… See the full description on the dataset page: https://huggingface.co/datasets/kth8/gemma-4-E4B-it-GPQA-Diamond-benchmark.gemma-4-E2B-it-MMLU-Pro-benchmarkBenchmark of google/gemma-4-E2B-it against MMLU-Pro dataset. Model's answer is considered correct if it matches the ground truth answer index exactly.
Accuracy: 61.6% with Python tool.
Metric
Value
Correct
617
Incorrect
381
Errors
3
Total samples
1001
Python tool calls
314
Python tool errors
12
Total completion tokens
1,499,382
gemma-4-E4B-it-ValleyBench-benchmarkBenchmark of google/gemma-4-E4B-it against kth8/ValleyBench dataset. Model's answer is considered correct if it is within 0.01 of ground answer.
Accuracy: 80.3% with Python tool.
Metric
Value
Correct
4014
Incorrect
958
Errors
28
Total samples
5000
Python tool calls
4843
Total completion tokens
4,595,312
Raw stats:
{
"accuracy": 0.803,
"correct": 4014,
"incorrect": 958,
"error": 28,
"total": 5000,
"python_tool_calls": 4843,
"completion_tokens": 4595312
}
