datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.DeepSeek-V4-Pro-Distilled-200K
DeepSeek‑V4‑Pro‑Distilled‑200K
High-quality Math & STEM reasoning distilled from DeepSeek‑V4‑Pro in Max mode
Reasoning traces · Proofs · Verification · Mathematics · Physics · Chemistry · Biology
Overview
DeepSeek‑V4‑Pro‑Distilled‑200K is a supervised fine-tuning collection of long-form mathematical and scientific reasoning. Its responses were generated with DeepSeek‑V4‑Pro in Max inference mode, then normalized into a compact conversational… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Pro-Distilled-200K.deepseek-v4-pro-agentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-agent.DeepSeek-V4-Distill-8000x
🐳 DeepSeek-V4-Distill-8100x
Dataset Summary
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash.
After the cleaning process, the released train split contains 7,716 high-quality JSONL examples.
[!NOTE]
The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Distill-8000x.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.Scale-SWE-Distilled-DeepSeek-v4-Pro-High-41k
Immersion in the GitHub Universe: Scaling Coding Agents to Mastery
🔥 Highlights
Source from 6M+ pull requests and 23000+ repositories.
Cover 5200 Repositories.
100k high-quality instances.
71k trajectories from DeepSeek v3.2 with 3.5B token.
Strong performance: 64% in SWE-bench-Verified trained from Qwen3-30A3B-Instruct.
📣 News
2026-06-11 We released DeNovoSWE, a high qaulity large scale long-horizon SWE dataset that is ready for… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/Scale-SWE-Distilled-DeepSeek-v4-Pro-High-41k.Deepseek-v4-pro-max-distill-1500x
Overeview
This dataset contains reasoning traces and final answers generated by DeepSeek-V4-Pro
(reasoning_effort=max, thinking.enabled=true) using prompts sampled from
Jackrong/GLM-5.1-Reasoning-1M-Cleaned.
Goal: just check quality
Update: The dataset have fully 1000 samples in 04/27/2026 cost only ~$5.46
Planning: try out another distill style such as roleplay
beyoru/Aesir-Character-CoT-roleplay
beyoru/Aisaka-LN-v4
Why DeepSeek-V4-Pro instead of OpenAI /… See the full description on the dataset page: https://huggingface.co/datasets/beyoru/Deepseek-v4-pro-max-distill-1500x.DeepSeek-V4.1-Flash-NVFP4-metrics
DeepSeek-V4.1-Flash-NVFP4 metrics
Everything behind the numbers in AtomicChat/DeepSeek-V4.1-Flash-NVFP4-nvidia.
logprobs/lp-<run>-<corpus>.npz: the raw top-512 log probabilities of every measurement run, 49,152 scored
positions each: ref, ref-repeat, ref-r3, ref-b1 (batch size 1) for the original; flat, flat-r2,
flat-r3 for the uncalibrated cast; nvidia, nvidia-r2, nvidia-r3 for the calibrated checkpoint.
logs/kld-<run>-<corpus>.json: the KL lower bound per run against ref… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/DeepSeek-V4.1-Flash-NVFP4-metrics.gdpval-deepseek-v4-proterminal-bench-2.1-deepseek-v4-flash-trajectories
DeepSeek-V4-Flash on Terminal-Bench 2.1 — full agent trajectories
Every agent trajectory from a controlled scaffold comparison: the same model, the
same machine, the same 89 tasks, only the agent harness changed.
Scaffold
Solved
terminus-2 (Terminal-Bench's own agent)
53 / 89
dsh sdk-minimal (DeepSeek Harness)
61 / 89
Paired: dsh solved 18 tasks terminus-2 missed, terminus-2 solved 8 that dsh
missed, 43 both, 18 neither, 2 not scorable (see Caveats).… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/terminal-bench-2.1-deepseek-v4-flash-trajectories.DeepSeek-V4-Flash-0731-REAM-calibration-stats
DeepSeek-V4-Flash-0731 — expert calibration statistics (REAM line)
Layerwise routed-expert statistics of
deepseek-ai/DeepSeek-V4-Flash-0731
(43 MoE layers × 256 experts), collected by running the full model over a
~4.9M-token multi-domain calibration mix (multi-turn dialogs, thinking and
direct modes, rendered with the model's own chat encoder). These are the
statistics behind the REAM144/96 release line — published so that expert
selection, pruning, merging and routing research… See the full description on the dataset page: https://huggingface.co/datasets/WaveCut/DeepSeek-V4-Flash-0731-REAM-calibration-stats.deepseek-v4-pro-agent-tool-calling-trajectory
DeepSeek V4 Pro ToolScale Agent SFT Dataset
A curated subset of multi-turn tool-calling trajectories generated by DeepSeek V4 Pro on ToolScale. The dataset is filtered by action-match score against ground-truth trajectories and is designed for supervised fine-tuning of agentic models on realistic, multi-step tool use.
Each conversation includes natural-language user requests, tool calls, tool observations, assistant reasoning traces, and grounded final responses across five… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/deepseek-v4-pro-agent-tool-calling-trajectory.deepseek-v4-tiny-cpu-repro-v1
DeepSeek-V4 tiny corrected-native-primitives CPU text fixture
Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class
using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction.
No upstream weights, paid GPU/cloud compute or useful-model claim.
This is not unmodified native Transformers or the complete production release.
Architecture and scope
Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.Wikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x
DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows
Teacher-distillation corpus generated with
deepseek-ai/DeepSeek-V4-Flash-0731.
The original manifest contained 45,000 unique seeds.
Following generation, QC, retry-based repair, quarantine auditing,
and recovery adjudication, 40,513 rows were retained.
Composition
Bucket
Rows
Coding
5,601
Agentic
9,982
Cyber blue
13,000
Controlled cyber red
6,999
Tool use
4,931
Total
40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.DeepSeek-v4-Flash-ChatThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Teich Test
This directory contains newline-delimited JSON training examples generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-flash.
Rows: 6313
Format
Each file is newline-delimited JSON where every line is already a training example.
Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Flash-Chat.deepseek-v4-flash-0731-m3-ultra
DeepSeek-V4-Flash-0731 on M3 Ultra 512 GB — benchmark dataset
Independent performance characterization of
Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX
on a single Mac Studio M3 Ultra (80-core GPU, 512 GB unified memory).
Engine: omlx 0.5.7 · OS: macOS 26.6 (25G72) · MLX: 0.32.0
Recommended configuration
omlx serve --model-dir /opt/models --port 8033 \
--hot-cache-max-size 256GB --initial-cache-blocks 512
// ~/.omlx/model_settings.json
{"version": 1, "models":… See the full description on the dataset page: https://huggingface.co/datasets/guruswami-ai/deepseek-v4-flash-0731-m3-ultra.scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928
Scale-SWE DeepSeek V4 Flash 0731 Think Rollouts
Successful AweAgent trajectories generated with deepseek-v4-flash-0731 in think mode.
Dataset summary
Source task instances: 3,393
Rollouts per source instance: 4
Total attempted rollouts: 13,572
Successful exported trajectories: 7,928
Unique instances represented by successful trajectories: 2,250
Scaffold: aweagent
Tool-call format: openai_function
The export retains assistant reasoning_content, function tool… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928.deepseek-v4-flash-rocm-vllm-repro
Reproducing DeepSeek-V4-Flash on AMD ROCm with vLLM: 32K Correctness and TopK Sweep
This article summarizes an engineering reproduction of
deepseek-ai/DeepSeek-V4-Flash on an AMD ROCm ModelScope DSW instance. The work
focuses on a practical question: can a complex, fast-moving DeepSeek-V4-Flash
serving path be turned into a reproducible ROCm baseline with explicit
correctness gates?
The answer from this run is yes, with an important boundary: the current setup
is a fallback-heavy… See the full description on the dataset page: https://huggingface.co/datasets/lyydfys/deepseek-v4-flash-rocm-vllm-repro.deepseek-v4-pro-max-distill-1k
Overeview
This dataset contains reasoning traces and final answers generated by DeepSeek-V4-Pro
(reasoning_effort=max, thinking.enabled=true) using prompts sampled from
Jackrong/GLM-5.1-Reasoning-1M-Cleaned.
Goal: just check quality
Update: The dataset have fully 1000 samples in 04/27/2026 cost only ~$5.46
Planning: try out another distill style such as roleplay
Why DeepSeek-V4-Pro instead of OpenAI / Anthropic?
For distillation, the teacher must expose… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-max-distill-1k.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.deepseek-v4-flash-reap-observations-v1
deepseek-v4-flash-reap-observations-v1
Observations/calibration artifacts collected from the deepseek-v4-flash lineage during REAP/quantization runs.
Provenance
Dataset provenance: contains model prompts, traces, and/or generated outputs collected during REAP/quantization runs. Output content follows the upstream model's usage terms, but no dataset license is asserted here. Contact the uploader before redistribution.
Acknowledgements
Cerebras… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/deepseek-v4-flash-reap-observations-v1.Deepseek-v4.1-CoTDatasets created by deepseek-ai/DeepSeek-V4.1-Flash.
You can use them to distill other models.
Ty!
DeepSeek-V4-Pro-Reasoning-8000x
DeepSeek-V4-Pro-Reasoning-8000x
This dataset contains 8,014 synthetic reasoning examples generated with DeepSeek V4 Pro through the DeepSeek API.
The release is branded as 8000x for readability, while the exact row count is 8,014.
This dataset is designed for supervised fine-tuning, reasoning distillation, and experimentation with long-form visible reasoning traces.
Dataset Summary
Release label: 8000x
Actual rows: 8,014
Teacher model: DeepSeek-V4-Pro… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Pro-Reasoning-8000x.Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule!
Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills:
Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability.
Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.DeepSeek_V4_Flash_distilled_dataset_5k
DeepSeek V4 Flash — Distilled Reasoning Dataset
A synthetic dataset of 5,099 unique reasoning traces designed to mirror the step-by-step thinking style of DeepSeek V4 Flash. Generated entirely with template-based parameterized generation (no LLM API calls).
Format
JSONL (one JSON object per line):
{
"id": "ds4f_math_000042",
"domain": "mathematics",
"subdomain": "algebra",
"difficulty": "easy",
"prompt": "Solve 3x + 7 = 22.",
"reasoning_trace":… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DeepSeek_V4_Flash_distilled_dataset_5k.deepseek-v4-flash-swebench-replay
deepseek-v4-flash-swebench-replay
中文
这是一个 DeepSeek V4 Flash 在 SWE-bench 上的 agentic replay 数据集仓库。
它的目标是让使用者不需要部署 SWE-bench,也不需要复现 Docker/benchmark 环境,就可以直接查看和重放模型的多轮推理与工具调用轨迹。
当前包含的数据
verified_agentic
lite_agentic
当前不包含的数据
单轮 single-turn trace
verified_mini_agentic(当前本地仅完成 31/50,因此不纳入首版)
分数汇总
verified_agentic: 354 / 500, Acc/Pass@1 = 70.8
lite_agentic: 182 / 300, Acc/Pass@1 = 60.67
数据来源
这些轨迹由… See the full description on the dataset page: https://huggingface.co/datasets/fxiao0369/deepseek-v4-flash-swebench-replay.
