mlfoundations-dev/Bespoke-Stratos-7B_1743996665_eval_0981
mlfoundations-dev/Bespoke-Stratos-7B_1743996665_eval_0981 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AIME25 AMC23 MATH500 GPQADiamond LiveCodeBench Accuracy 16.0 14.7 57.5 77.0 31.3 27.3 AIME24 Average Accuracy: 16.00% ± 1.74% Number of Runs: 5 Run Accuracy Questions Solved Total Questions 1 13.33% 4 30 2 10.00% 3 30 3 16.67% 5 30 4 20.00% 6 30 5 20.00% 6 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Bespoke-Stratos-7B_1743996665_eval_0981.
06
mlfoundations-dev/Bespoke-Stratos-7B1743996665eval_0981
Precomputed model outputs for evaluation.
Evaluation Results
Summary
AIME24
- Average Accuracy: 16.00% ± 1.74%
- Number of Runs: 5
AIME25
- Average Accuracy: 14.67% ± 0.73%
- Number of Runs: 5
AMC23
- Average Accuracy: 57.50% ± 1.22%
- Number of Runs: 5
MATH500
- Accuracy: 77.00% | Accuracy | Questions Solved | Total Questions | |----------|-----------------|----------------| | 77.00% | 385 | 500 |
GPQADiamond
- Average Accuracy: 31.31% ± 0.71%
- Number of Runs: 3
LiveCodeBench
- Average Accuracy: 27.33% ± 1.03%
- Number of Runs: 3
