datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jailbreak-deepseek-v3.2-expDeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.aime_1983_2023_deepseek-r1_traces_16384AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
DeepSeek-R1-Distill-Qwen-7B_eval_d81a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MMLUPro
HMMT
HLE
AIME25
LiveCodeBenchv5
Accuracy
43.4
25.0
12.4
36.0
34.5
MMLUPro
Accuracy: 43.38%
Accuracy
Questions Solved
Total Questions
43.38%
N/A
N/A
HMMT
Average Accuracy: 25.00% ± 1.72%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.deepseek-leetcodeDeepseek Leetcode dataset from https://github.com/deepseek-ai/DeepSeek-Coder/tree/main/Evaluation/LeetCode
deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.Chinese-DeepSeek-R1-Distill-data-110k
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。
该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.aime_1983_2023_deepseek-r1_traces_32768DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AIME25
AMC23
GPQADiamond
MATH500
Accuracy
42.7
22.7
67.0
33.3
79.6
AIME24
Average Accuracy: 42.67% ± 4.75%
Number of Runs: 5
Run
Accuracy
Questions Solved
Total Questions
1
50.00%
15
30
2
26.67%
8
30
3
53.33%
16
30
4
50.00%
15
30
5
33.33%
10
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.DeepSeek-R1-Distill-Qwen-7B_eval_118b
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 31.18% ± nan%
Number of Runs: 1
Run
Accuracy
Questions Solved
Total Questions
1
31.18%
87
279
aime_1983_2023_deepseek-r1-distill-qwen-14b_traces_32768aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_32768deepseek-hermes-reasoning-traces
DeepSeek V4 Pro Hermes Reasoning Traces
19,331 multi-turn ChatML + Hermes reasoning traces generated by DeepSeek V4 Pro. Designed for LoRA fine-tuning local models to operate as Hermes Agent instances.
Quick Start
\
Splits
Split
Traces
train
16,431
valid
1,933
test
967
Variants (VRAM-Tiered)
Variant
Max Tokens
Traces
GPU
nano
2,048
15,948
Dev / 7B
budget
4,096
2,149
48GB
standard
8,192
990
64GB
spark
16,384
244… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-hermes-reasoning-traces.details_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private
Dataset Card for Evaluation run of hosted_vllm//fsx/anton/deepseek-r1-checkpoint
Dataset automatically created during the evaluation run of model hosted_vllm//fsx/anton/deepseek-r1-checkpoint.
The dataset is composed of 15 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_hosted_vllm____fsx__anton__deepseek-r1-checkpoint_private.DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
32.7
71.8
80.8
31.1
32.5
31.1
27.2
8.8
8.5
15.0
15.3
23.7
15.4
AIME24
Average Accuracy: 32.67% ± 2.39%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554.Chinese-DeepSeek-R1-Distill-data-110k-SFT
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下:
Math:共计36568个样本,
Exam:共计2432个样本,
STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.aime_1983_2023_deepseek-r1-distill-qwen-1.5b_traces_32768eval-DeepSeek-R1-0528
r1-0528 Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.865
math_pass@1:64_samples
64
0.0%
aime25
0.831
math_pass@1:64_samples
64
0.0%
arenahard
0.951
eval/overall_winrate
500
0.0%
bbh_generative
0.894
extractive_match
1
0.0%
creative-writing-v3
0.803
creative_writing_score
96
0.0%
drop_generative_nous
0.865
drop_acc
1
0.0%
eqbench3
0.865
eqbench_score
135
0.0%
gpqa_diamond
0.781
gpqa_pass@1:8_samples8
0.1%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-R1-0528.deepseek-v4-pro-agent-tool-calling-trajectory
DeepSeek V4 Pro ToolScale Agent SFT Dataset
A curated subset of multi-turn tool-calling trajectories generated by DeepSeek V4 Pro on ToolScale. The dataset is filtered by action-match score against ground-truth trajectories and is designed for supervised fine-tuning of agentic models on realistic, multi-step tool use.
Each conversation includes natural-language user requests, tool calls, tool observations, assistant reasoning traces, and grounded final responses across five… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/deepseek-v4-pro-agent-tool-calling-trajectory.DeepSeek-1.5B_mmlu-pro_16384_train0test8DeepSeek-R1-Distill-Qwen-7B_eval_c64a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_c64a
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_v3
Average Accuracy: 30.47% ± 0.69%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
31.34%
84
268
2
30.97%
83
268
3
29.10%
78
268
eval-DeepSeek-V3-0324
dsv3 Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.506
math_pass@1:64_samples
64
100.0%
aime25
0.422
math_pass@1:64_samples
64
100.0%
arenahard
0.926
eval/overall_winrate
500
0.0%
bbh_generative
0.868
extractive_match
1
100.0%
creative-writing-v3
0.767
creative_writing_score
96
0.0%
drop_generative_nous
0.829
drop_acc
1
100.0%
eqbench3
0.831
eqbench_score
135
0.0%
gpqa_diamond
0.680
gpqa_pass@1:8_samples8
100.0%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-V3-0324.deepseek-v4-tiny-cpu-repro-v1
DeepSeek-V4 tiny corrected-native-primitives CPU text fixture
Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class
using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction.
No upstream weights, paid GPU/cloud compute or useful-model claim.
This is not unmodified native Transformers or the complete production release.
Architecture and scope
Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.Wikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_16384DeepSeek-R1-Distill-Llama-8B-max-activation-SAE-cache-L7Created using https://github.com/KoyenaPal/autointerp/blob/master/demo/cache.py
Datasets: cerebras/SlimPajama-627B and koyena/OpenR1-Math-220k-formatted
SAE: https://huggingface.co/fnlp/Llama-Scope-R1-Distill/tree/main/400M-Slimpajama-400M-OpenR1-Math-220k/L7R
