kth8/gemma-4-E4B-it-SuperGPQA-benchmark
Benchmark of google/gemma-4-E4B-it against m-a-p/SuperGPQA dataset. Accuracy: 38.1% with Python tool. Metric Value Correct 761 Incorrect 1239 Errors 0 Total samples 2000 Python tool calls 200 Python tool errors 10 Total completion tokens 4,253,773 Raw stats: { "accuracy": 0.381, "correct": 761, "incorrect": 1239, "error": 0, "total": 2000, "python_tool_calls": 200, "python_tool_errors": 10, "completion_tokens": 4253773 }
license: apache-2.0 language:
- en base_model: google/gemma-4-E4B-it datasets:
- m-a-p/SuperGPQA --- Benchmark of google/gemma-4-E4B-it against m-a-p/SuperGPQA dataset.
Accuracy: 38.1% with Python tool. | Metric | Value | |----------------------|---------------| | Correct | 761 | | Incorrect | 1239 | | Errors | 0 | | Total samples | 2000 | | Python tool calls| 200 | | Python tool errors| 10 | | Total completion tokens | 4,253,773 |
Raw stats:
{
"accuracy": 0.381,
"correct": 761,
"incorrect": 1239,
"error": 0,
"total": 2000,
"python_tool_calls": 200,
"python_tool_errors": 10,
"completion_tokens": 4253773
}