kth8/Qwen3.6-27B-AWQ-BF16-INT4-SuperGPQA-benchmark
Benchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset. Accuracy: 69.2% with Python tool. Metric Value Correct 692 Incorrect 295 Errors 13 Total samples 1000 Python tool calls 1508 Total completion tokens 3,806,045 Raw stats: { "accuracy": 0.692, "correct": 692, "incorrect": 295, "error": 13, "total": 1000, "python_tool_calls": 1508, "completion_tokens": 3806045 }
license: apache-2.0 language:
- en base_model: cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 datasets:
- m-a-p/SuperGPQA --- Benchmark of cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 against m-a-p/SuperGPQA dataset.
Accuracy: 69.2% with Python tool. | Metric | Value | |----------------------|---------------| | Correct | 692 | | Incorrect | 295 | | Errors | 13 | | Total samples | 1000 | | Python tool calls| 1508 | | Total completion tokens | 3,806,045 |
Raw stats:
{
"accuracy": 0.692,
"correct": 692,
"incorrect": 295,
"error": 13,
"total": 1000,
"python_tool_calls": 1508,
"completion_tokens": 3806045
}