datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phi_24K_qwq_6K_eval_2e29
mlfoundations-dev/phi_24K_qwq_6K_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
35.0
66.2
80.8
8.0
41.9
46.6
1.2
3.2
2.0
27.7
3.6
0.5
AIME24
Average Accuracy: 35.00% ± 2.22%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
36.67%
11
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_24K_qwq_6K_eval_2e29.phi-2-labeled
Dataset Card for "phi-1"
More Information needed
phi-2-embeddings
Dataset Card for "phi-2-embeddings"
More Information needed
phi2-spl-evaluationphi_27K_qwq_3K_eval_2e29
mlfoundations-dev/phi_27K_qwq_3K_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
31.3
57.0
63.4
7.4
42.6
47.1
1.2
0.5
0.1
23.7
2.9
0.6
AIME24
Average Accuracy: 31.33% ± 1.26%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
30.00%
9
30
2
26.67%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/phi_27K_qwq_3K_eval_2e29.
