datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepSeek-V3-Infinity-Instruct-0625
DeepSeek-V3-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside DeepSeek-V3-0324 as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with deepseek-ai/DeepSeek-V3-0324 at temperature=1.
For more details on the training methodology and… See the full description on the dataset page: https://huggingface.co/datasets/nebius/DeepSeek-V3-Infinity-Instruct-0625.Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High
Distill
This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format.
Dataset Structure
The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.Superpotion-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Superpotion-DeepSeek-V3.2.Speciale is a dataset containing structured medical reasoning responses, testing the limits of DeepSeek V3.2 Speciale's medical reasoning skills across a wide variety of medical disciplines and tasks!
This dataset contains:
28.8k synthetically generated medical prompts, with all responses generated using DeepSeek V3.2 Speciale.
Structured medical reasoning: Superpotion uses organized, informative… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Superpotion-DeepSeek-V3.2-Speciale.Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k-formatted
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k-formatted
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k
Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
このデータセットはNSFW表現を含みます。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(R-18)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k.Raiden-Mini-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Raiden-Mini-DeepSeek-V3.2.Speciale is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek-V3.2.Speciale's reasoning skills!
This dataset contains:
a default subset of ~8k 'creative_content' and 'analytical_reasoning' prompts from sequelbox/Raiden-DeepSeek-R1, with all responses generated by DeepSeek V3.2 Speciale.
provides an unfiltered look into the reasoning skills of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-Mini-DeepSeek-V3.2-Speciale.UML-Generator-Dataset-DeepSeek-V3.2Click here to support our open-source dataset and model releases!
UML-Generator-Dataset-DeepSeek-V3.2 is a dataset focused on analysis and code-reasoning, creating UML diagrams testing the limits of DeepSeek V3.2's modeling and design skills!
This dataset contains:
2.7k synthetically generated prompts to create UML diagrams in response to user input, with all responses generated using DeepSeek V3.2.
All responses contain a multi-step thinking process to perform effective analysis, followed by… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/UML-Generator-Dataset-DeepSeek-V3.2.Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k
Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。
データの詳細
各データは以下のキーを含んでいます。
major_genre: ジャンル(大分類)
minor_genre: ジャンル(小分類)
tag: 年齢制限用タグ(全年齢、R-15)
world_setting: 舞台・世界観の設定
scene_setting: 対話シーンの設定
user_setting: ユーザー側のキャラクターの設定
assistant_setting: アシスタント側のキャラクターの設定
dialogue_tone: 対話のトーン
conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)
設定等の情報からsystem… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k.Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted
Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20kにsystem messageを追加して整形したデータセットです。
データの詳細については元データセットのREADMEを参照してください。
ライセンス
MITライセンスの元配布します。
Superpotion-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases!
Superpotion-DeepSeek-V3.2.Speciale is a dataset containing structured medical reasoning responses, testing the limits of DeepSeek V3.2 Speciale's medical reasoning skills across a wide variety of medical disciplines and tasks!
This dataset contains:
28.8k synthetically generated medical prompts, with all responses generated using DeepSeek V3.2 Speciale.
Structured medical reasoning: Superpotion uses organized, informative… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/Superpotion-DeepSeek-V3.2-Speciale.Titanium3-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases!
Titanium3-DeepSeek-V3.1-Terminus is a dataset focused on architecture and DevOps, testing the limits of DeepSeek V3.1 Terminus's architect and coding skills!
This dataset contains:
27.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek V3.1 Terminus in reasoning mode:
20k selected technical expertise prompts from sequelbox/Titanium2.1-DeepSeek-R1 focused on… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Titanium3-DeepSeek-V3.1-Terminus.DES-Reasoning-DeepSeek-V3.1Click here to support our open-source dataset and model releases!
DES-Reasoning-DeepSeek-V3.1 is a dataset focused on analysis and reasoning, creating discrete event simulations testing the limits of DeepSeek V3.1's simulation, Python scripting, and analysis skills!
This dataset contains:
4.03k synthetically generated prompts to create discrete event simulations and analysis chat in response to user input, with all responses generated using DeepSeek V3.1.
All responses contain a multi-step… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DES-Reasoning-DeepSeek-V3.1.qs-deepseek-platinum-v3
QS DeepSeek Platinum v3
High-quality training dataset for DeepSeek trading model.
Dataset Details
Total samples: 17,212
Format: OpenAI messages format (system/user/assistant)
Sources:
qs-deepseek-platinum-15k (HF)
Synthetic negative examples
Live signals from Supabase
Data Quality
Metric
Value
Positive PnL
74.4%
Negative PnL
25.5%
Avg message length
2,220 chars
Max message length
8,922 chars
Label Distribution… See the full description on the dataset page: https://huggingface.co/datasets/henryzhang2024/qs-deepseek-platinum-v3.Titanium3-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases!
Titanium3-DeepSeek-V3.1-Terminus is a dataset focused on architecture and DevOps, testing the limits of DeepSeek V3.1 Terminus's architect and coding skills!
This dataset contains:
27.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek V3.1 Terminus in reasoning mode:
20k selected technical expertise prompts from sequelbox/Titanium2.1-DeepSeek-R1 focused on… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium3-DeepSeek-V3.1-Terminus.DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.orca-dpo-deepseek-v3.2
Orca DPO - DeepSeek V3.2
An updated successor to argilla/distilabel-intel-orca-dpo-pairs, bringing the preference pairs from the GPT-4 era into 2026.
We took the original Intel/orca_dpo_pairs prompts, generated fresh responses with DeepSeek V3.2, and scored all pairs with Skywork Reward V2 to determine which response is preferred.
What's in the dataset
Each row contains a prompt with two responses — a winner (chosen) and a loser (rejected) — determined by reward model… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/orca-dpo-deepseek-v3.2.Tachibana3-Part1-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases!
Tachibana3-Part1-DeepSeek-V3.1-Terminus is a dataset focused on high-difficulty code production tasks, testing the limits of DeepSeek V3.1 Terminus's code-reasoning skills!
This dataset contains 9.3k high-difficulty code-production prompts:
Questions prioritize real-world, challenging coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, mobile, gamedev… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana3-Part1-DeepSeek-V3.1-Terminus.Tachibana3-Part2-DeepSeek-V3.2Click here to support our open-source dataset and model releases!
Tachibana3-Part2-DeepSeek-V3.2 is a dataset focused on high-difficulty code production tasks, testing the limits of DeepSeek V3.2's code-reasoning skills!
This dataset contains 9.3k high-difficulty code-production prompts:
Questions prioritize real-world, challenging coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, mobile, gamedev, cloud, QA, custom… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana3-Part2-DeepSeek-V3.2.DeepSeek-V3_Synthetic_Conversation_Dialogue
DeepSeek V3 Synthetic Conversation Dialogue Dataset
This dataset contains basic conversational dialogue for a chatbot with system prompts.
Source
DeepSeek V3 Model was used to generate this synthetic dataset.
deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
deepseek-v3.2-SWE-Agent
DeepSeek V3.2 SWE-Agent Execution Traces
Full execution traces of DeepSeek V3.2 running SWE-Agent on SWE-Bench.
Dataset Structure
Each JSON file in traces/ corresponds to one SWE-Bench problem instance. The filename is the instance ID (e.g., django__django-12345.json).
Trace Schema
{
"instance_id": "django__django-12345",
"model": "DeepSeek-V3.2",
"agent": "SWE-agent",
"total_steps": 15,
"total_run_duration_seconds": 120.5,
"exit_status":… See the full description on the dataset page: https://huggingface.co/datasets/zanderjiang/deepseek-v3.2-SWE-Agent.kanitakorn-deepseek-v39-unicode-micro
Kanitakorn DeepSeek v39 Unicode Micro
Small repaired ThaiExam-style SFT mix for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 540 rows = 500 MCQ + 40 identity
MCQ label balance: a=100 b=100 c=100 d=100 e=100
Audit: readable UTF-8 Thai, no mojibake markers, valid final-answer format
Constraints:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v39-unicode-micro.
