datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jailbreak-deepseek-v3.2-explygo-public-witness-feed
LYGO Public Witness — public RESOURCE overlay
Snapshot of public HTTPS GET sources (USGS, EONET, ISS, NWS, GDACS, NOAA SWPC, Open-Meteo, OpenSky sample, launches, CoinGecko, Celestrak, lattice mirrors).
Every point is RESOURCE
Failed sources stay named SHADOW (no invented coordinates)
Not private intelligence. Not a live Star Chart write.
Live globe: https://chatagent.ca/witness/
excavationpro-music-stream
Excavationpro public music stream (160 kbps)
Owner / artist: Justin Helmer · Excavationpro · LightfatherPolicy: Own-work only. Public discovery streams (not DistroKid-dependent).Lattice signature: Δ9Φ963-PUBLIC-MUSIC-STREAM-v1
Listen
https://deepseekoracle.github.io/Excavationpro/excavationpro-listen.html
http://asiancoastline.com/ (custom domain music portal)
Layout
Path
Role
stream/<sha256>.mp3
Flat 160k streams (~first 10k −… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/excavationpro-music-stream.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.aime_1983_2023_deepseek-r1_traces_16384AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository.
AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of
high-quality and challenging reasoning problems. These problems are collected from numerous
open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination.
All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone
rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.train_v6_deepseekAM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.EvolInstruct_zh_DeepseekAPI和之前的Evol-Instruction尝试对比(https://huggingface.co/datasets/lorinma/Chinese_Evol_Instruct_3.5),使用了中文prompt。
因为OpenAI接口太贵,使用了DeepSeek赠送的1000万token。这次生成了一万条基本用完了。
一共有3个文件:
combined_seed_correct.json 是使用的基础种子任务371条,alpaca格式。使用了 Belle的中文种子任务175条。并且参照了 4 增加了ShareGPT的数据以更接近真实世界的用法,掺入了 Wildchat-zh抽样196条 ,多轮对话只采用第一个有意义的问答对。
evolve_chinese.py 基于H2O EvolInstruction的代码。
0227_evol_combinedseedcorrect.json 生成的1.2万条数据。
APIGen-MT-5k-with-cot-v1-deepseek_deepseekdeepseek-v2-codder-minecraft-apimmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
lygo-protocol-stack
LYGO Protocol Stack — Hugging Face mirror
Canonical GitHub: DeepSeekOracle/lygo-protocol-stack
Contents
P0 Nano Kernel — Φ-gate (AMPLIFY / SOFTEN / QUARANTINE), 42 hardened test vectors, Python/C/Rust parity
P1–P5 — Memory Mycelium, Cognitive Bridge, Vortex Consensus, Ascension Engine, Harmony Node
ClawHub catalog — links + skills.json (full skill trees on GitHub)
P0 quick demo (local)
pip install pytest
python tools/run_p0_demo.py
python… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/lygo-protocol-stack.DeepSeek-R1-Distill-Qwen-7B_eval_d81a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MMLUPro
HMMT
HLE
AIME25
LiveCodeBenchv5
Accuracy
43.4
25.0
12.4
36.0
34.5
MMLUPro
Accuracy: 43.38%
Accuracy
Questions Solved
Total Questions
43.38%
N/A
N/A
HMMT
Average Accuracy: 25.00% ± 1.72%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.deepseek-leetcodeDeepseek Leetcode dataset from https://github.com/deepseek-ai/DeepSeek-Coder/tree/main/Evaluation/LeetCode
emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it
Per-token emotion trajectories, DeepSeek-written stories
google/gemma-4-31b-it read over the DeepSeek-written three-emotion stories. Pairs with the Gemma-written set to separate what the model does from what the story writer does.
Each story is stored as one .npz. The arrays are per token, so a trajectory
can be replayed word by word rather than only summarised.
Contents
Path
Contents
shards/<story_id>.npz
one story, arrays below
manifest.jsonl
one… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it.numina_amc_aime_deepseek_r1_responsesmath7500_train_DeepSeek-R1-Distill-Qwen-1p5_32K_tokensAIME_Deepseek_CleanDeepSeek_Math_V2AM-DeepSeek-R1-0528-Distilled
📘 Dataset Summary
This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher.
A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.DeepSeek-V4-Pro-Distilled-200K
DeepSeek‑V4‑Pro‑Distilled‑200K
High-quality Math & STEM reasoning distilled from DeepSeek‑V4‑Pro in Max mode
Reasoning traces · Proofs · Verification · Mathematics · Physics · Chemistry · Biology
Overview
DeepSeek‑V4‑Pro‑Distilled‑200K is a supervised fine-tuning collection of long-form mathematical and scientific reasoning. Its responses were generated with DeepSeek‑V4‑Pro in Max inference mode, then normalized into a compact conversational… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Pro-Distilled-200K.logiqa-deepseek-v3synthetic-medical-conversations-deepseek-v3
🍎 Synthetic Multipersona Doctor Patient Conversations.
Author: Nisten Tahiraj
License: MIT
🧠 Generated by DeepSeek V3 running in full BF16.
🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival systems for reducing hallucinations and increasing the diagnosis quality.
🐧 Conversations… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/synthetic-medical-conversations-deepseek-v3.deepseek-v4-pro-agentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-agent.DeepSeek-V4-Distill-8000x
🐳 DeepSeek-V4-Distill-8100x
Dataset Summary
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash.
After the cleaning process, the released train split contains 7,716 high-quality JSONL examples.
[!NOTE]
The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Distill-8000x.Chinese-DeepSeek-R1-Distill-data-110k
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。
该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.DeepSeek-Prover-V1
Evaluation Results |
Model & Dataset Downloads |
License |
Contact
Paper Link👁️
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
1. Introduction
Proof assistants like Lean have revolutionized mathematical proof verification, ensuring high accuracy and reliability. Although large language models (LLMs) show promise in… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-Prover-V1.
