datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime_1983_2023_deepseek-r1_traces_16384mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
DeepSeek-R1-Distill-Qwen-7B_eval_d81a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MMLUPro
HMMT
HLE
AIME25
LiveCodeBenchv5
Accuracy
43.4
25.0
12.4
36.0
34.5
MMLUPro
Accuracy: 43.38%
Accuracy
Questions Solved
Total Questions
43.38%
N/A
N/A
HMMT
Average Accuracy: 25.00% ± 1.72%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.aime_1983_2023_deepseek-r1_traces_32768DeepSeek-R1-Distill-Qwen-7B_eval_118b
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 31.18% ± nan%
Number of Runs: 1
Run
Accuracy
Questions Solved
Total Questions
1
31.18%
87
279
DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AIME25
AMC23
GPQADiamond
MATH500
Accuracy
42.7
22.7
67.0
33.3
79.6
AIME24
Average Accuracy: 42.67% ± 4.75%
Number of Runs: 5
Run
Accuracy
Questions Solved
Total Questions
1
50.00%
15
30
2
26.67%
8
30
3
53.33%
16
30
4
50.00%
15
30
5
33.33%
10
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981.aime_1983_2023_deepseek-r1-distill-qwen-14b_traces_32768DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
32.7
71.8
80.8
31.1
32.5
31.1
27.2
8.8
8.5
15.0
15.3
23.7
15.4
AIME24
Average Accuracy: 32.67% ± 2.39%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554.aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_32768aime_1983_2023_deepseek-r1-distill-qwen-1.5b_traces_32768DeepSeek-R1-Distill-Qwen-7B_eval_c64a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_c64a
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_v3
Average Accuracy: 30.47% ± 0.69%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
31.34%
84
268
2
30.97%
83
268
3
29.10%
78
268
DeepSeek-R1-Distill-Llama-8B-max-activation-SAE-cache-L7Created using https://github.com/KoyenaPal/autointerp/blob/master/demo/cache.py
Datasets: cerebras/SlimPajama-627B and koyena/OpenR1-Math-220k-formatted
SAE: https://huggingface.co/fnlp/Llama-Scope-R1-Distill/tree/main/400M-Slimpajama-400M-OpenR1-Math-220k/L7R
aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_16384DeepSeek-R1-Distill-Qwen-1.5B-Self-CalibrationThis dataset contains data for the paper Efficient Test-Time Scaling via Self-Calibration.
We propose an efficient test-time scaling method by using model confidence for dynamically sampling adjustment, since confidence can be seen as an intrinsic measure that directly reflects model uncertainty on different tasks. For example, we can incorporate the model’s confidence into self-consistency by assigning each sampled response $y_i$ a confidence score $c_i$. Instead of treating all responses… See the full description on the dataset page: https://huggingface.co/datasets/HINT-lab/DeepSeek-R1-Distill-Qwen-1.5B-Self-Calibration.details_deepseek-ai__DeepSeek-R1-Distill-Llama-70B
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Llama-70B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Llama-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deepseek-ai__DeepSeek-R1-Distill-Llama-70B.deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
Dataset Card for deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel/raw/main/pipeline.yaml"… See the full description on the dataset page: https://huggingface.co/datasets/dmitriihook/deepseek-r1-qwen-32b-planning-6-blocks-self-probing-state-distilabel.DeepSeek-R1-Distill-Qwen-7B_eval_d54a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d54a
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5total
Average Accuracy: 43.33% ± 0.20%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
43.64%
384
880
2
43.41%
382
880
3
42.95%
378
880
DeepSeek-R1-Distill-Qwen-7B_OpenThoughts3_eval_8179
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_OpenThoughts3_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
64.0
91.2
89.0
65.4
47.3
61.8
23.3
24.2
51.3
10.9
45.5
35.3
AIME24
Average Accuracy: 64.00% ± 1.23%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_OpenThoughts3_eval_8179.aime_1983_2023_deepseek-r1-distill-qwen-14b_traces_16384details_nbeerbower__DeepSeek-R1-Qwen-lorablated-32B
Dataset Card for Evaluation run of nbeerbower/DeepSeek-R1-Qwen-lorablated-32B
Dataset automatically created during the evaluation run of model nbeerbower/DeepSeek-R1-Qwen-lorablated-32B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_nbeerbower__DeepSeek-R1-Qwen-lorablated-32B.details_huihui-ai__DeepSeek-R1-Distill-Qwen-32B-abliterated
Dataset Card for Evaluation run of huihui-ai/DeepSeek-R1-Distill-Qwen-32B-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/DeepSeek-R1-Distill-Qwen-32B-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_huihui-ai__DeepSeek-R1-Distill-Qwen-32B-abliterated.DeepSeek-R1-Distill-Llama-8B-max-activation-SAE-cache-L23Chinese-DeepSeek-R1-Distill-data-110k-decontaminated
Decontaminated — Congliu/Chinese-DeepSeek-R1-Distill-data-110k
What this is
A filtered version of Congliu/Chinese-DeepSeek-R1-Distill-data-110k (revision
8520b649430617c2be4490f424d251d09d835ed3) with exact-duplicate rows and rows overlapping standard benchmark test sets
removed. This is a different artifact from the companion contamination report — that one is an
audit of what's wrong; this one is the corpus with those rows actually taken out, ready to train on.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/Chinese-DeepSeek-R1-Distill-data-110k-decontaminated.math500-cot-deepseek-r1-1.5b
MATH-500 CoT completions (DeepSeek-R1-Distill-Qwen-1.5B)
Successful chain-of-thought completions for HuggingFaceH4/MATH-500 test problems, generated with deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B via vLLM.
Files
File
Description
records.parquet
Main dataset: correct completions as token IDs
manifest.json
Schema, tokenizer, run ids, decoding config
problem_index.json
unique_id → problem_idx in MATH-500 test
subject_max_tokens.json
Per-subject completion… See the full description on the dataset page: https://huggingface.co/datasets/ChrisMcCormick/math500-cot-deepseek-r1-1.5b.openr1-clean-DeepSeek-R1-Distill-Qwen-7B-generations
16 generations per prompt from deepseek distilled qwen 7b for openr1
dataset_info:
features:
- name: message_id
dtype: int64
- name: message_response_id
dtype: string
- name: messages
list:
- name: content
dtype: string
- name: role
dtype: string
- name: response
dtype: string
- name: response_id
dtype: int64
- name: verified_answer
dtype: string
- name: extracted_answer
dtype: string
- name: is_correct… See the full description on the dataset page: https://huggingface.co/datasets/wen-sun/openr1-clean-DeepSeek-R1-Distill-Qwen-7B-generations.DeepSeek-R1-20k
Rethinking Generalization in Reasoning SFT
This repository contains datasets associated with the paper "Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability".
The research investigates the factors influencing cross-domain generalization in Large Language Models (LLMs) during reasoning-focused supervised fine-tuning (SFT) with long chain-of-thought (CoT) data.
Key Findings
Optimization Dynamics: Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/jasonrqh/DeepSeek-R1-20k.PlugNPlayComp-STAR-41K-DPOMixedPolicy-DeepSeek-R1-Distill-Qwen-1.5BPlugNPlayComp-STAR-41K-DPOMixedPolicy-DeepSeek-R1-Distill-Qwen-7Bmath7500_train_solutions_DeepSeek-R1-Distill-Qwen-7B_32K_tokensPlugNPlayComp-STAR-41K-MixedPolicy-DeepSeek-R1-Distill-Qwen-1.5B
