datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kimik25-fp4-vllm-isl8192osl1024conc128DeepSeek-R1-FP4_1757573629_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757573629_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 61.11% ± 0.48%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
61.11%
121
198
2
60.10%
119
198
3
62.12%
123
198
DeepSeek-R1-FP4_1757550665_eval_d91d
chengfu0118/DeepSeek-R1-FP4_1757550665_eval_d91d
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 88.60%
Accuracy
Questions Solved
Total Questions
88.60%
443
500
DeepSeek-R1-FP4_1757557889_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757557889_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 61.28% ± 0.77%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
63.13%
125
198
2
60.61%
120
198
3
60.10%
119
198
dsr1-fp4-sgl-isl8192osl1024DeepSeek-R1-FP4_1757548772_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757548772_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 59.26% ± 1.07%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
61.62%
122
198
2
57.07%
113
198
3
59.09%
117
198
DeepSeek-R1-FP4_1757540347_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757540347_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 16.67% ± 0.63%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
17.17%
34
198
2
17.68%
35
198
3
15.15%
30
198
DeepSeek-R1-FP4_1757570920_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757570920_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 59.93% ± 0.77%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
60.61%
120
198
2
61.11%
121
198
3
58.08%
115
198
dsr1-fp4-sgl-isl1024osl8192DeepSeek-R1-FP4_1757536719_eval_2870
chengfu0118/DeepSeek-R1-FP4_1757536719_eval_2870
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 68.33% ± 1.18%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%
21
30
2
63.33%
19
30
3
66.67%
20
30
4
63.33%
19
30
5
70.00%
21
30
6
70.00%
21
30
7
66.67%
20
30
8
70.00%
21
30
9
66.67%
20
30
10
76.67%
23
30
DeepSeek-R1-FP4_1757537053_eval_2870
chengfu0118/DeepSeek-R1-FP4_1757537053_eval_2870
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 65.00% ± 2.27%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
73.33%
22
30
2
70.00%
21
30
3
70.00%
21
30
4
56.67%
17
30
5
56.67%
17
30
6
73.33%
22
30
7
66.67%
20
30
8
70.00%
21
30
9
53.33%
16
30
10
60.00%
18
30
DeepSeek-R1-FP4_1757547837_eval_d91d
chengfu0118/DeepSeek-R1-FP4_1757547837_eval_d91d
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 88.20%
Accuracy
Questions Solved
Total Questions
88.20%
441
500
DeepSeek-R1-FP4_1757561992_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757561992_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 60.44% ± 1.91%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
64.65%
128
198
2
56.57%
112
198
3
60.10%
119
198
DeepSeek-R1-FP4_1757543283_eval_7c1d
chengfu0118/DeepSeek-R1-FP4_1757543283_eval_7c1d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
MATH500
GPQADiamond
Accuracy
0.7
15.0
7.4
AIME24
Average Accuracy: 0.67% ± 0.63%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
30
2
0.00%
0
30
3
0.00%
0
30
4
0.00%
0
30
5
6.67%
2
30
6
0.00%
0
30
7
0.00%
0
30
8
0.00%
0
30
9
0.00%
0
30
10
0.00%
0… See the full description on the dataset page: https://huggingface.co/datasets/chengfu0118/DeepSeek-R1-FP4_1757543283_eval_7c1d.gptoss-fp4-vllm-isl1024osl1024gptoss-fp4-vllm-isl8192osl1024kimik25-fp4-vllm-isl8192osl1024conc32DeepSeek-R1-FP4_1757536509_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757536509_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 61.78% ± 0.36%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
62.63%
124
198
2
61.62%
122
198
3
61.11%
121
198
DeepSeek-R1-FP4_1757539348_eval_d91d
chengfu0118/DeepSeek-R1-FP4_1757539348_eval_d91d
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 88.00%
Accuracy
Questions Solved
Total Questions
88.00%
440
500
DeepSeek-R1-FP4_1757542177_eval_d91d
chengfu0118/DeepSeek-R1-FP4_1757542177_eval_d91d
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 89.00%
Accuracy
Questions Solved
Total Questions
89.00%
445
500
DeepSeek-R1-FP4_1757541768_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757541768_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 63.13% ± 0.71%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
64.65%
128
198
2
61.62%
122
198
3
63.13%
125
198
DeepSeek-R1-FP4_1757541727_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757541727_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 60.10% ± 1.33%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
63.13%
125
198
2
59.60%
118
198
3
57.58%
114
198
DeepSeek-R1-FP4_1757545008_eval_d91d
chengfu0118/DeepSeek-R1-FP4_1757545008_eval_d91d
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 87.60%
Accuracy
Questions Solved
Total Questions
87.60%
438
500
DeepSeek-R1-FP4_1757545324_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757545324_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 60.44% ± 0.69%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
62.12%
123
198
2
59.60%
118
198
3
59.60%
118
198
DeepSeek-R1-FP4_1757545947_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757545947_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 62.63% ± 0.24%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
62.12%
123
198
2
62.63%
124
198
3
63.13%
125
198
DeepSeek-R1-FP4_1757548999_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757548999_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 63.13% ± 1.09%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
63.64%
126
198
2
60.61%
120
198
3
65.15%
129
198
DeepSeek-R1-FP4_1757550047_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757550047_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 59.93% ± 1.40%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
56.57%
112
198
2
61.11%
121
198
3
62.12%
123
198
DeepSeek-R1-FP4_1757552390_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757552390_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 60.77% ± 1.69%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
64.65%
128
198
2
57.58%
114
198
3
60.10%
119
198
DeepSeek-R1-FP4_1757553968_eval_f912
chengfu0118/DeepSeek-R1-FP4_1757553968_eval_f912
Precomputed model outputs for evaluation.
Evaluation Results
GPQADiamond
Average Accuracy: 61.62% ± 0.86%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
62.12%
123
198
2
59.60%
118
198
3
63.13%
125
198
DeepSeek-R1-FP4_1757559098_eval_d91d
chengfu0118/DeepSeek-R1-FP4_1757559098_eval_d91d
Precomputed model outputs for evaluation.
Evaluation Results
MATH500
Accuracy: 86.80%
Accuracy
Questions Solved
Total Questions
86.80%
434
500
