datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.Titanium4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Titanium 4 is an agentic coding dataset focused on DevOps and architecture, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture, incident response, configuration and cost optimization, security… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro.Mitakihara2-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Mitakihara 2 is an agentic coding dataset focused on MLOps and AI development, testing the limits of DeepSeek-V4-Pro's agentic skills:
Questions prioritize real-world, challenging agentic coding tasks in AI development, research, deployment, interpretability, operation and experimentation. The primary purpose of the Mitakihara dataset series is to accelerate and decentralize AI… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro.deepseek-v4-pro-tachibana4Click here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-tachibana4.Tachibana4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule!
This is an early sneak preview of Tachibana 4, containing the first 1.2k rows!
Tachibana 4 is an upcoming agentic coding dataset, generated by DeepSeek-V4-Pro:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics.
Areas of focus include back-end and front-end development, systems programming, distributed systems… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro-PREVIEW.Titanium4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule!
This is an early sneak preview of Titanium 4, containing the first 4.9k rows!
Titanium 4 is an upcoming agentic coding dataset focused on DevOps and architecture, generated by DeepSeek-V4-Pro:
Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics.
Areas of focus include IaC, cloud architecture… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro-PREVIEW.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.deepseek-v4-reasoning-code-2500
Mirror: lucsaint/Deepseek-V4-Reasoning-Code-2500
Pinned snapshot / mirror of lucsaint/Deepseek-V4-Reasoning-Code-2500, re-hosted for PROTISEC
research reproducibility. Redistributed under the upstream license (apache-2.0)
with attribution — all credit to the original author.
Original author: lucsaint
Source dataset: lucsaint/Deepseek-V4-Reasoning-Code-2500
License: apache-2.0
Family: coding_traces
Mode: full
Rows cached: 2556
Changes vs upstream: cached snapshot, possibly… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-reasoning-code-2500.
