datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qs-deepseek-platinum-v3
QS DeepSeek Platinum v3
High-quality training dataset for DeepSeek trading model.
Dataset Details
Total samples: 17,212
Format: OpenAI messages format (system/user/assistant)
Sources:
qs-deepseek-platinum-15k (HF)
Synthetic negative examples
Live signals from Supabase
Data Quality
Metric
Value
Positive PnL
74.4%
Negative PnL
25.5%
Avg message length
2,220 chars
Max message length
8,922 chars
Label Distribution… See the full description on the dataset page: https://huggingface.co/datasets/henryzhang2024/qs-deepseek-platinum-v3.DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.DeepSeek-V3_Synthetic_Conversation_Dialogue
DeepSeek V3 Synthetic Conversation Dialogue Dataset
This dataset contains basic conversational dialogue for a chatbot with system prompts.
Source
DeepSeek V3 Model was used to generate this synthetic dataset.
deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript.
kanitakorn-deepseek-v39-unicode-micro
Kanitakorn DeepSeek v39 Unicode Micro
Small repaired ThaiExam-style SFT mix for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 540 rows = 500 MCQ + 40 identity
MCQ label balance: a=100 b=100 c=100 d=100 e=100
Audit: readable UTF-8 Thai, no mojibake markers, valid final-answer format
Constraints:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v39-unicode-micro.
