datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
full-nla-aa-step10001.00-eps6full-nla-aa-step10001-eps5type134_step1000tmp07type12_plus_sftloss_step1000tmp10bokeh-eval-infer-step10000-kesctype12_plus_sftloss_step1000tmp07ot3_300k_more_ckpts_step1000_eval_27e9
mlfoundations-dev/ot3_300k_more_ckpts_step1000_eval_27e9
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBench
Average Accuracy: 34.64% ± 0.87%
Number of Runs: 6
Run
Accuracy
Questions Solved
Total Questions
1
33.86%
173
511
2
30.92%
158
511
3
36.79%
188
511
4
35.62%
182
511
5
34.44%
176
511
6
36.20%
185
511
type134_step1000tmp10
