datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
a1_code_sql_create_context_1744691318_eval_1331
mlfoundations-dev/a1_code_sql_create_context_1744691318_eval_1331
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
GPQADiamond
JEEBench
MMLUPro
LiveCodeBench
CodeElo
Accuracy
9.7
40.0
52.0
26.9
20.4
28.4
11.7
4.3
AIME24
Average Accuracy: 9.67% ± 1.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
6.67%
2
30
2
16.67%
5
30
3
16.67%
5
30
4
6.67%
2
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_sql_create_context_1744691318_eval_1331.sql-context-instructions
Dataset Card for "sql-context-instructions"
More Information needed
a1_code_sql_create_context_eval_636d
mlfoundations-dev/a1_code_sql_create_context_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
9.0
40.0
53.2
27.8
20.6
28.8
12.1
3.6
5.3
AIME24
Average Accuracy: 9.00% ± 0.67%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
6.67%
2
30
2
10.00%
3
30
3
6.67%
2
30
4
10.00%
3… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/a1_code_sql_create_context_eval_636d.
