datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
s1K-step-conditional-control-old
Citation Information
@misc{muennighoff2025s1simpletesttimescaling,
title={s1: Simple test-time scaling},
author={Niklas Muennighoff and Zitong Yang and Weijia Shi and Xiang Lisa Li and Li Fei-Fei and Hannaneh Hajishirzi and Luke Zettlemoyer and Percy Liang and Emmanuel Candès and Tatsunori Hashimoto},
year={2025},
eprint={2501.19393},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2501.19393},
}
s1K-1.1-tfidf-sweep-kmeans-dim10000-20251118s1_s1k_difficult_questionsreasoning-s1K-1.1-noxmls1K-1.1-modernbert-split-kmeans-dim384-20251118reasoning-s1K-1.1-evaluation-presynths1K-1.1-modernbert-split-kmeans-dim768-20250916reasoning-s1K-1.1-evaluationqwen3_sft_s1k_saving_onlyreasoning-s1K-1.1-evaluation-noxmlHuggingFaceH4__R2-Q7B-GR1-ALL-s1k-5e-5-weight-decay-1e-4_privatereasoning-s1K-1.1noxmlreasoning-s1K-1.1bunnycore__Qwen-2.5-7b-S1k-details
Dataset Card for Evaluation run of bunnycore/Qwen-2.5-7b-S1k
Dataset automatically created during the evaluation run of model bunnycore/Qwen-2.5-7b-S1k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen-2.5-7b-S1k-details.s1_s1k_minwait_tokenized_wait0s1K-modernbert-split-kmeans-dim768-20250321s1k-1.1-test-192_eval_2870
mlfoundations-dev/s1k-1.1-test-192_eval_2870
Precomputed model outputs for evaluation.
Evaluation Results
AIME24
Average Accuracy: 10.33% ± 1.45%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
10.00%
3
30
2
13.33%
4
30
3
6.67%
2
30
4
0.00%
0
30
5
13.33%
4
30
6
16.67%
5
30
7
6.67%
2
30
8
10.00%
3
30
9
13.33%
4
30
10
13.33%
4
30
s1k_1k_samples_1024_thinking_20251120_230635s1K-modernbert-split-kmeans-dim768-20250211bunnycore__Maestro-S1k-7B-Sce-details
Dataset Card for Evaluation run of bunnycore/Maestro-S1k-7B-Sce
Dataset automatically created during the evaluation run of model bunnycore/Maestro-S1k-7B-Sce
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Maestro-S1k-7B-Sce-details.reasoning-s1K-1.1-synths1K-sharegpt_1743202401_eval_0771
mlfoundations-dev/s1K-sharegpt_1743202401_eval_0771
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AIME25
MATH500
Accuracy
14.0
8.7
54.4
AIME24
Average Accuracy: 14.00% ± 1.12%
Number of Runs: 5
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2
16.67%
5
30
3
13.33%
4
30
4
16.67%
5
30
5
10.00%
3
30
AIME25
Average Accuracy: 8.67% ± 1.52%
Number of… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/s1K-sharegpt_1743202401_eval_0771.b2_math_fasttext_pos_s1k_reformat_neg_lap1official_mathb2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_10k_eval_636d
mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_10k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
22.3
60.5
80.0
28.6
39.9
37.5
22.7
6.1
8.9
AIME24
Average Accuracy: 22.33% ± 1.70%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
23.33%
7
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_10k_eval_636d.b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_1k_eval_636d
mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_1k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
14.7
54.8
76.6
28.6
38.2
36.0
22.4
4.2
6.0
AIME24
Average Accuracy: 14.67% ± 1.43%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
13.33%
4
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_1k_eval_636d.b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_0.3k_eval_636d
mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_0.3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
17.7
53.0
74.6
29.2
39.7
39.6
25.4
4.2
5.5
AIME24
Average Accuracy: 17.67% ± 0.82%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
16.67%
5
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_0.3k_eval_636d.b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_3k_eval_636d
mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_3k_eval_636d
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
Accuracy
19.0
58.8
79.4
29.6
41.5
38.7
22.7
4.6
5.2
AIME24
Average Accuracy: 19.00% ± 1.25%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
26.67%
8
30
2… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_3k_eval_636d.b2_math_fasttext_pos_s1k_reformat_neg_lap1official_math_10ks1K-1.1-minilm-split-kmeans-dim384-20251118s1k_modified
