datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime_1983_2023_deepseek-r1_traces_16384mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
DeepSeek-R1-Distill-Qwen-7B_eval_d81a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MMLUPro
HMMT
HLE
AIME25
LiveCodeBenchv5
Accuracy
43.4
25.0
12.4
36.0
34.5
MMLUPro
Accuracy: 43.38%
Accuracy
Questions Solved
Total Questions
43.38%
N/A
N/A
HMMT
Average Accuracy: 25.00% ± 1.72%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.numina_amc_aime_deepseek_r1_responsesMagpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B.aime_1983_2023_deepseek-r1_traces_32768DeepSeek-R1-Distill-Qwen-7B_eval_118b
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 31.18% ± nan%
Number of Runs: 1
Run
Accuracy
Questions Solved
Total Questions
1
31.18%
87
279
DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AIME25
AMC23
GPQADiamond
MATH500
Accuracy
42.7
22.7
67.0
33.3
79.6
AIME24
Average Accuracy: 42.67% ± 4.75%
Number of Runs: 5
Run
Accuracy
Questions Solved
Total Questions
1
50.00%
15
30
2
26.67%
8
30
3
53.33%
16
30
4
50.00%
15
30
5
33.33%
10
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981.AM-DeepSeek-R1-Distilled-1.4Maime_1983_2023_deepseek-r1-distill-qwen-14b_traces_32768DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
32.7
71.8
80.8
31.1
32.5
31.1
27.2
8.8
8.5
15.0
15.3
23.7
15.4
AIME24
Average Accuracy: 32.67% ± 2.39%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554.science_traces_original_DeepSeek-R1-Distill-Qwen-32Baime_1983_2023_deepseek-r1-distill-qwen-7b_traces_32768aime_1983_2023_deepseek-r1-distill-qwen-1.5b_traces_32768filtered_science_DeepSeek-R1-Distill-Qwen-32B_judged_science_traces_original_DeepSeek-R1-Distill-Qwen-32BDeepSeek-R1-Distill-Qwen-7B_eval_c64a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_c64a
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_v3
Average Accuracy: 30.47% ± 0.69%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
31.34%
84
268
2
30.97%
83
268
3
29.10%
78
268
PARTIAL-STAR41K-DeepSeek-R1-Distill-Qwen-7B-Size-16-Blockwise-InterIntra-Attention-0_8192deepseek-reasoner-full-supergpqa-r1DeepSeek-R1-only-CoTdetails_mobiuslabsgmbh__DeepSeek-R1-ReDistill-Llama3-8B-v1.1_v2
Dataset Card for Evaluation run of mobiuslabsgmbh/DeepSeek-R1-ReDistill-Llama3-8B-v1.1
Dataset automatically created during the evaluation run of model mobiuslabsgmbh/DeepSeek-R1-ReDistill-Llama3-8B-v1.1.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_mobiuslabsgmbh__DeepSeek-R1-ReDistill-Llama3-8B-v1.1_v2.AM-DeepSeek-R1-0528-Distilled-with-Systemdetails_mobiuslabsgmbh__DeepSeek-R1-ReDistill-Qwen-7B-v1.1_v2
Dataset Card for Evaluation run of mobiuslabsgmbh/DeepSeek-R1-ReDistill-Qwen-7B-v1.1
Dataset automatically created during the evaluation run of model mobiuslabsgmbh/DeepSeek-R1-ReDistill-Qwen-7B-v1.1.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_mobiuslabsgmbh__DeepSeek-R1-ReDistill-Qwen-7B-v1.1_v2.Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B-formatteddetails_deepseek-ai__DeepSeek-R1-Distill-Qwen-14B_v2
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deepseek-ai__DeepSeek-R1-Distill-Qwen-14B_v2.DeepSeek-R1-Distill-Llama-8B-max-activation-SAE-cache-L7Created using https://github.com/KoyenaPal/autointerp/blob/master/demo/cache.py
Datasets: cerebras/SlimPajama-627B and koyena/OpenR1-Math-220k-formatted
SAE: https://huggingface.co/fnlp/Llama-Scope-R1-Distill/tree/main/400M-Slimpajama-400M-OpenR1-Math-220k/L7R
details_deepseek-ai__DeepSeek-R1-Distill-Qwen-32B_v2
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deepseek-ai__DeepSeek-R1-Distill-Qwen-32B_v2.numina_math_deepseek_r1_responsesaime_1983_2023_deepseek-r1-distill-qwen-7b_traces_16384DeepSeek-R1-Distill-Qwen-1.5B-Self-CalibrationThis dataset contains data for the paper Efficient Test-Time Scaling via Self-Calibration.
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model uncertainty on different tasks. For example, we can incorporate the model’s confidence into self-consistency by assigning each sampled response $y_i$ a confidence score $c_i$. Instead of treating all responses… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/DeepSeek-R1-Distill-Qwen-1.5B-Self-Calibration.
