deepseekR1
DeepSeek-R1-Distill-Qwen-7B_eval_d81a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MMLUPro
HMMT
HLE
AIME25
LiveCodeBenchv5
Accuracy
43.4
25.0
12.4
36.0
34.5
MMLUPro
Accuracy: 43.38%
Accuracy
Questions Solved
Total Questions
43.38%
N/A
N/A
HMMT
Average Accuracy: 25.00% ± 1.72%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.DeepSeek-R1-Distill-Qwen-7B_eval_118b
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 31.18% ± nan%
Number of Runs: 1
Run
Accuracy
Questions Solved
Total Questions
1
31.18%
87
279
DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AIME25
AMC23
GPQADiamond
MATH500
Accuracy
42.7
22.7
67.0
33.3
79.6
AIME24
Average Accuracy: 42.67% ± 4.75%
Number of Runs: 5
Run
Accuracy
Questions Solved
Total Questions
1
50.00%
15
30
2
26.67%
8
30
3
53.33%
16
30
4
50.00%
15
30
5
33.33%
10
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981.DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
32.7
71.8
80.8
31.1
32.5
31.1
27.2
8.8
8.5
15.0
15.3
23.7
15.4
AIME24
Average Accuracy: 32.67% ± 2.39%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554.deepseek-r1-distill-cotCreated by several different models:
DeepSeek R1
DeepSeek R1 Distill Qwen 14B
Qwen3.8 27B
format:
{"quest": ..., "aswer": "..."}
ty
DeepSeek-R1-Distill-Qwen-7B_eval_c64a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_c64a
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_v3
Average Accuracy: 30.47% ± 0.69%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
31.34%
84
268
2
30.97%
83
268
3
29.10%
78
268
