datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Terminal-Lego-Traj-Deepseek-V3-2-15kdeepseek-v3.2-speciale-openr1-math-3kInspired by @OpenR1
The questions for this dataset were all sourced from the first 3.3k prompts in open-r1/OpenR1-Math-220k
Dataset Stats (provided by OpenRouter):
Cost: $ 21.1 (USD)
Tokens (input + output): 52.3 M
deepseek-v3.2-speciale-OpenCodeReasoning-3kThe questions for this dataset were all sourced from the first 3k prompts in nvidia/OpenCodeReasoning
Dataset Stats (provided by OpenRouter):
Cost: $ 19.2 (USD)
Tokens (input + output): 47 M
Chinese-DeepSeek-V3.2-Exp-chat-example
deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本
一、前言
本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。
二、数据与方法
数据来源:用户构建的 6,655 轮真实中文对话样本。
估算方法:
中文字符近似为 1 Token;
英文 4 字符 ≈ 1 Token;
用于规模与上下文预算对比,而非精确 Token 计数。
统计维度:
平均 Prompt/Output 长度(字符与估算 Token);
总 Token 占上下文窗口比例;
语言分布(Prompt 语言类型);
对话长度分布(用户提问、助手回答、总对话长度)。
三、总体结果
1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.qs-deepseek-platinum-v3
QS DeepSeek Platinum v3
High-quality training dataset for DeepSeek trading model.
Dataset Details
Total samples: 17,212
Format: OpenAI messages format (system/user/assistant)
Sources:
qs-deepseek-platinum-15k (HF)
Synthetic negative examples
Live signals from Supabase
Data Quality
Metric
Value
Positive PnL
74.4%
Negative PnL
25.5%
Avg message length
2,220 chars
Max message length
8,922 chars
Label Distribution… See the full description on the dataset page: https://huggingface.co/datasets/henryzhang2024/qs-deepseek-platinum-v3.DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.deepseek-v3.2-speciale-1000xThis is a reasoning dataset created using Deepseek v3.2 Speciale with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Deepseek v3.2 Speciale by fine-tuning already existing open-source LLMs.
Stats
Costs: $ 1.46 (USD)
Total tokens (input + output): 3.61 M
DeepSeek-V3-Distill-Cybersecurity-en
🔐 CyberSecPentest-DS: Deepseek-V3 Distilled Dataset
📂 Dataset Overview:This is a high-quality distilled dataset 🧪, specialized in the cybersecurity penetration testing domain, generated using Deepseek-V3 🚀. It provides curated knowledge for red-teaming, vulnerability assessment, and ethical hacking research.
💡 Key Features:
🛡️ Focus: Penetration testing, exploit development, and security assessments.
🧠 Source: Knowledge distilled from Deepseek-V3 for accuracy &… See the full description on the dataset page: https://huggingface.co/datasets/Bouquets/DeepSeek-V3-Distill-Cybersecurity-en.kira-wiki-instruct_deepseek-v3-0324DeepSeek-V3.2-Exp-reasoning-example
🐳 DeepSeek-V3.2-Exp-reasoning vs DeepSeek-R1-0528: Math Reasoning Comparison 🍎
Note: DeepSeek-R1-0528 has no explicit chain-of-thought, while deepseek-ai/DeepSeek-V3.2-Exp (abbrev. V3.2-Exp) produces answers with structured derivations. This report was analyzed by GPT-5-Extended-Thinking. The sample size is small; conclusions are for reference only.
Author: Soren
1. Executive Summary
Sample size: 208 problems (mixed types).
Average steps (reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V3.2-Exp-reasoning-example.Deepseek-V3-Distilled-Ancient-Chinese-Translation
Dataset Card
这是一个文言文/白话文互译的高质量数据集,翻译精准,理由充分。总共有70K条数据,通过Deepseek-V3蒸馏获得。
Dataset Card Authors
Shen ZhuoKang From ECNU
Dataset Card Contact
10235101553@stu.ecnu.edu.cn
deepseek-v3.2-speciale-OpenCodeReasoning-3kThe questions for this dataset were all sourced from the first 3k prompts in nvidia/OpenCodeReasoning
Dataset Stats (provided by OpenRouter):
Cost: $ 19.2 (USD)
Tokens (input + output): 47 M
DeepSeek-V3_Synthetic_Conversation_Dialogue
DeepSeek V3 Synthetic Conversation Dialogue Dataset
This dataset contains basic conversational dialogue for a chatbot with system prompts.
Source
DeepSeek V3 Model was used to generate this synthetic dataset.
Orion-Deepseek-V3-RP-Filtereddeepseek-v3.1-1000xThis is a reasoning dataset created using DeepSeek V3.1. Some of these questions are from reedmayhew and the rest were generated.
Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting.
The dataset is meant for creating distilled versions of DeepSeek V3.1 by fine-tuning already existing open-source LLMs.
deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
deepseek-v3.1-200xThis is a reasoning dataset created using DeepSeek V3.1. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of DeepSeek V3.1 by fine-tuning already existing open-source LLMs.
claude-4.5-opus-high-reasoning-X-DeepseekV3.2-264xThis is a reasoning dataset created using Claude Opus 4.5 with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Claude Opus 4.5 by fine-tuning already existing open-source LLMs.
Stats
Costs: $ 52.3 (USD)
Total tokens (input + output): 2.13 M
kira-wiki-instruct-GenQuery_deepseek-v3-0324deepseek-v3-distill-RWA-1000
A Practical Exploration of Mixed-Style Response LLMs via Few-Shot LoRA Fine-Tuning
1. Research Background and Motivation
With the increasingly widespread application of Large Language Models (LLMs) today, how to make model outputs more transparent, natural, and understandable has become an important research direction. Traditional LLMs typically output the final answer directly, making their internal reasoning process a "black box" to the user. To enhance the… See the full description on the dataset page: https://huggingface.co/datasets/dianzinao/deepseek-v3-distill-RWA-1000.kanitakorn-deepseek-v39-unicode-micro
Kanitakorn DeepSeek v39 Unicode Micro
Small repaired ThaiExam-style SFT mix for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 540 rows = 500 MCQ + 40 identity
MCQ label balance: a=100 b=100 c=100 d=100 e=100
Audit: readable UTF-8 Thai, no mojibake markers, valid final-answer format
Constraints:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v39-unicode-micro.TE_dataset_deepseekv3_2.jsonlDeepseek-V3-RPdeepseek-v3-distill-RWA-10000deepseek-v3neo-synthkinkgweep gwoop
health_deepseekV3
