CoolFace
Datasetpublic

kth8/Qwen3.5-4B-Claude-Opus-Reasoning-Distill-SuperGPQA-benchmark

Benchmark of TeichAI/Qwen3.5-4B-Claude-Opus-Reasoning-Distill against m-a-p/SuperGPQA dataset. Accuracy: 41.5% with Python tool. Metric Value Correct 415 Incorrect 573 Errors 11 Total samples 999 Python tool calls 2527 Total completion tokens 4,149,159 Raw stats: { "accuracy": 0.415, "correct": 415, "incorrect": 573, "error": 11, "total": 999, "python_tool_calls": 2527, "completion_tokens": 4149159 }

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
0likes27downloads
Dataset Card

license: apache-2.0 language:

Accuracy: 41.5% with Python tool. | Metric | Value | |----------------------|---------------| | Correct | 415 | | Incorrect | 573 | | Errors | 11 | | Total samples | 999 | | Python tool calls| 2527 | | Total completion tokens | 4,149,159 |

Raw stats:

json
{
  "accuracy": 0.415,
  "correct": 415,
  "incorrect": 573,
  "error": 11,
  "total": 999,
  "python_tool_calls": 2527,
  "completion_tokens": 4149159
}