CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lvogel123 /jailbreak-deepseek-v3.2-exptabular1K<n<10K1 likes16k downloads11mo agoHugging Face02DeepSeekOracle /lygo-public-witness-feed LYGO Public Witness — public RESOURCE overlay Snapshot of public HTTPS GET sources (USGS, EONET, ISS, NWS, GDACS, NOAA SWPC, Open-Meteo, OpenSky sample, launches, CoinGecko, Celestrak, lattice mirrors). Every point is RESOURCE Failed sources stay named SHADOW (no invented coordinates) Not private intelligence. Not a live Star Chart write. Live globe: https://chatagent.ca/witness/ 1 likes11k downloads10m agoHugging Face03DeepSeekOracle /excavationpro-music-stream Excavationpro public music stream (160 kbps) Owner / artist: Justin Helmer · Excavationpro · LightfatherPolicy: Own-work only. Public discovery streams (not DistroKid-dependent).Lattice signature: Δ9Φ963-PUBLIC-MUSIC-STREAM-v1 Listen https://deepseekoracle.github.io/Excavationpro/excavationpro-listen.html http://asiancoastline.com/ (custom domain music portal) Layout Path Role stream/<sha256>.mp3 Flat 160k streams (~first 10k −… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/excavationpro-music-stream.audioaudio-to-audio10K<n<100K0 likes10k downloads7d agoHugging Face04TeichAI /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K106 likes5.4k downloads4mo agoHugging Face05jonathanyin /aime_1983_2023_deepseek-r1_traces_16384tabularn<1K0 likes2.9k downloads1y agoHugging Face06a-m-team /AM-DeepSeek-R1-Distilled-1.4MFor more open-source datasets, models, and methodologies, please visit our GitHub repository. AM-DeepSeek-R1-Distilled-1.4M is a large-scale general reasoning task dataset composed of high-quality and challenging reasoning problems. These problems are collected from numerous open-source datasets, semantically deduplicated, and cleaned to eliminate test set contamination. All responses in the dataset are distilled from the reasoning model (mostly DeepSeek-R1) and have undergone rigorous… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-Distilled-1.4M.text-generation1M<n<10M184 likes2.2k downloads1y agoHugging Face07prefixsliding /train_v6_deepseektext100K<n<1M0 likes2.1k downloads1y agoHugging Face08a-m-team /AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training. Due to certain constraints, we are only able to open-source a subset of the complete dataset. Model Training Performance based on our complete dataset On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.tabulartext-generation10M<n<100M56 likes2.1k downloads1y agoHugging Face09lorinma /EvolInstruct_zh_DeepseekAPI和之前的Evol-Instruction尝试对比(https://huggingface.co/datasets/lorinma/Chinese_Evol_Instruct_3.5),使用了中文prompt。 因为OpenAI接口太贵,使用了DeepSeek赠送的1000万token。这次生成了一万条基本用完了。 一共有3个文件: combined_seed_correct.json 是使用的基础种子任务371条,alpaca格式。使用了 Belle的中文种子任务175条。并且参照了 4 增加了ShareGPT的数据以更接近真实世界的用法,掺入了 Wildchat-zh抽样196条 ,多轮对话只采用第一个有意义的问答对。 evolve_chinese.py 基于H2O EvolInstruction的代码。 0227_evol_combinedseedcorrect.json 生成的1.2万条数据。 0 likes2k downloads3y agoHugging Face10dipta007 /APIGen-MT-5k-with-cot-v1-deepseek_deepseektext1K<n<10K0 likes1.9k downloads1y agoHugging Face11SkillFi /deepseek-v2-codder-minecraft-apitexttext-generationn<1K0 likes1.9k downloads1y agoHugging Face12Mumon /mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples. tabular10K<n<100K1 likes1.8k downloads2y agoHugging Face13DeepSeekOracle /lygo-protocol-stack LYGO Protocol Stack — Hugging Face mirror Canonical GitHub: DeepSeekOracle/lygo-protocol-stack Contents P0 Nano Kernel — Φ-gate (AMPLIFY / SOFTEN / QUARANTINE), 42 hardened test vectors, Python/C/Rust parity P1–P5 — Memory Mycelium, Cognitive Bridge, Vortex Consensus, Ascension Engine, Harmony Node ClawHub catalog — links + skills.json (full skill trees on GitHub) P0 quick demo (local) pip install pytest python tools/run_p0_demo.py python… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/lygo-protocol-stack.audio0 likes1.8k downloads22d agoHugging Face14mlfoundations-dev /DeepSeek-R1-Distill-Qwen-7B_eval_d81a mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a Precomputed model outputs for evaluation. Evaluation Results Summary Metric MMLUPro HMMT HLE AIME25 LiveCodeBenchv5 Accuracy 43.4 25.0 12.4 36.0 34.5 MMLUPro Accuracy: 43.38% Accuracy Questions Solved Total Questions 43.38% N/A N/A HMMT Average Accuracy: 25.00% ± 1.72% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.tabular10K<n<100K0 likes1.4k downloads1y agoHugging Face15ronaldcmz /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K0 likes1.3k downloads3mo agoHugging Face16r0b0tlab /deepseek-v4-pro-0813-agentic DeepSeek-V4-Pro 0813 Agentic (DS4) A standalone, verifiable-first agentic training corpus: 19,072 training traces plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families, each row admitted only after passing a deterministic programmatic verifier. The corpus is designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL (verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.tabulartext-generation10K<n<100K22 likes1.2k downloads1mo agoHugging Face17davidheineman /deepseek-leetcodeDeepseek Leetcode dataset from https://github.com/deepseek-ai/DeepSeek-Coder/tree/main/Evaluation/LeetCode tabularn<1K0 likes1.2k downloads1y agoHugging Face18abotresol /emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it Per-token emotion trajectories, DeepSeek-written stories google/gemma-4-31b-it read over the DeepSeek-written three-emotion stories. Pairs with the Gemma-written set to separate what the model does from what the story writer does. Each story is stored as one .npz. The arrays are per token, so a trajectory can be replayed word by word rather than only summarised. Contents Path Contents shards/<story_id>.npz one story, arrays below manifest.jsonl one… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it.feature-extraction0 likes1.1k downloads2mo agoHugging Face19hbXNov /numina_amc_aime_deepseek_r1_responsestextn<1K0 likes1.1k downloads2y agoHugging Face20nishadsinghi /math7500_train_DeepSeek-R1-Distill-Qwen-1p5_32K_tokens0 likes1.1k downloads2y agoHugging Face21Rock23210 /AIME_Deepseek_Cleantextn<1K0 likes1k downloads2y agoHugging Face22memmywinks /DeepSeek_Math_V20 likes1k downloads2mo agoHugging Face23a-m-team /AM-DeepSeek-R1-0528-Distilled 📘 Dataset Summary This dataset is a high-quality reasoning corpus distilled from DeepSeek-R1-0528, an improved version of the DeepSeek-R1 large language model. Compared to its initial release, DeepSeek-R1-0528 demonstrates significant advances in reasoning, instruction following, and multi-turn dialogue. Motivated by these improvements, we collected and distilled a diverse set of 2.6 million queries across multiple domains, using DeepSeek-R1-0528 as the teacher. A notable… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-R1-0528-Distilled.text-generation1M<n<10M102 likes1k downloads1y agoHugging Face24Jackrong /DeepSeek-V4-Pro-Distilled-200K DeepSeek‑V4‑Pro‑Distilled‑200K High-quality Math & STEM reasoning distilled from DeepSeek‑V4‑Pro in Max mode Reasoning traces · Proofs · Verification · Mathematics · Physics · Chemistry · Biology Overview DeepSeek‑V4‑Pro‑Distilled‑200K is a supervised fine-tuning collection of long-form mathematical and scientific reasoning. Its responses were generated with DeepSeek‑V4‑Pro in Max inference mode, then normalized into a compact conversational… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Pro-Distilled-200K.texttext-generation100K<n<1M10 likes1k downloads2mo agoHugging Face25fireworks-ai /logiqa-deepseek-v3text1K<n<10K0 likes1k downloads2y agoHugging Face26OnDeviceMedNotes /synthetic-medical-conversations-deepseek-v3 🍎 Synthetic Multipersona Doctor Patient Conversations. Author: Nisten Tahiraj License: MIT 🧠 Generated by DeepSeek V3 running in full BF16. 🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival systems for reducing hallucinations and increasing the diagnosis quality. 🐧 Conversations… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/synthetic-medical-conversations-deepseek-v3.text35 likes938 downloads2y agoHugging Face27ansulev /deepseek-v4-pro-agentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-agent.2 likes936 downloads4mo agoHugging Face28Jackrong /DeepSeek-V4-Distill-8000x 🐳 DeepSeek-V4-Distill-8100x Dataset Summary DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash. After the cleaning process, the released train split contains 7,716 high-quality JSONL examples. [!NOTE] The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Distill-8000x.texttext-generation1K<n<10K92 likes927 downloads5mo agoHugging Face29Congliu /Chinese-DeepSeek-R1-Distill-data-110k 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face&nbsp;&nbsp; | &nbsp;&nbsp;🤖 ModelScope &nbsp;&nbsp; | &nbsp;&nbsp;🚀 Github &nbsp;&nbsp; | &nbsp;&nbsp;📑 Blog 注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。 该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.tabulartext-generation100K<n<1M789 likes918 downloads2y agoHugging Face30deepseek-ai /DeepSeek-Prover-V1 Evaluation Results | Model & Dataset Downloads | License | Contact Paper Link👁️ DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 1. Introduction Proof assistants like Lean have revolutionized mathematical proof verification, ensuring high accuracy and reliability. Although large language models (LLMs) show promise in… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-Prover-V1.text10K<n<100K74 likes900 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.