CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sequelbox /Raiden-DeepSeek-R1Click here to support our open-source dataset and model releases! Raiden-DeepSeek-R1 is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek R1's reasoning skills! This dataset contains: 63k 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1, with all responses generated by deepseek-ai/DeepSeek-R1. Responses demonstrate the reasoning capabilities of DeepSeek's 685b parameter R1 reasoning model.… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-DeepSeek-R1.texttext-generation10K<n<100K53 likes189 downloads2y agoHugging Face02CreitinGameplays /DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1tabular1K<n<10K0 likes162 downloads2y agoHugging Face03sequelbox /Tachibana4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule! Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills: Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability. Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro.texttext-generation10K<n<100K21 likes145 downloads5mo agoHugging Face04Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes142 downloads4mo agoHugging Face05sequelbox /Titanium2-DeepSeek-R1Click here to support our open-source dataset and model releases! Titanium2-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills! This dataset contains: 32.4k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium2-DeepSeek-R1.texttext-generation10K<n<100K4 likes125 downloads1y agoHugging Face06Trelis /openassistant-deepseek-coder Chat Fine-tuning Dataset - OpenAssistant DeepSeek Coder This dataset allows for fine-tuning chat models using: B_INST = '\n### Instruction:\n' E_INST = '\n### Response:\n' BOS = '<|begin▁of▁sentence|>' EOS = '\n<|EOT|>\n' Sample Preparation: The dataset is cloned from TimDettmers, which itself is a subset of the Open Assistant dataset, which you can find here. This subset of the data only contains the highest-rated paths in the conversation tree, with a total of 9,846 samples. The… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/openassistant-deepseek-coder.text10K<n<100K10 likes121 downloads3y agoHugging Face07sequelbox /Celestia3-DeepSeek-R1-0528Click here to support our open-source dataset and model releases! Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1 0528's science-reasoning skills! This dataset contains: 90.9k synthetically generated science prompts, with all responses generated using DeepSeek R1 0528. Primary subjects are physics, chemistry, biology, and computer science; secondary subjects include Earth science, astronomy, and information theory. All prompts are synthetic, taken… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528.texttext-generation10K<n<100K35 likes121 downloads1y agoHugging Face08chatdeepai /deepseek-1m-context-benchmark DeepSeek 1M Context Benchmark This dataset is the publication-safe measurement release for DeepSeek 1M Context Benchmark: Retrieval Accuracy, Latency, and Cost, version v1.0.0. It contains 344 sanitized terminal API records produced by the frozen protocol deepseek-v4-long-context-retrieval-v1.1.0 during a bounded run from 2026-08-06T20:17:02.706Z through 2026-08-07T00:07:44.737Z. The study compared deepseek-v4-flash and deepseek-v4-pro on deterministic synthetic English… See the full description on the dataset page: https://huggingface.co/datasets/chatdeepai/deepseek-1m-context-benchmark.tabular1K<n<10K0 likes114 downloads18d agoHugging Face09sequelbox /Titanium4-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule! Titanium 4 is an agentic coding dataset focused on DevOps and architecture, testing the limits of DeepSeek-V4-Pro's agentic skills: Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics. Areas of focus include IaC, cloud architecture, incident response, configuration and cost optimization, security… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro.texttext-generation10K<n<100K10 likes103 downloads3mo agoHugging Face10sequelbox /Mitakihara2-DeepSeek-V4-ProClick here to support our open-source dataset and model releases - help us speed up our release schedule! Mitakihara 2 is an agentic coding dataset focused on MLOps and AI development, testing the limits of DeepSeek-V4-Pro's agentic skills: Questions prioritize real-world, challenging agentic coding tasks in AI development, research, deployment, interpretability, operation and experimentation. The primary purpose of the Mitakihara dataset series is to accelerate and decentralize AI… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara2-DeepSeek-V4-Pro.texttext-generation10K<n<100K4 likes101 downloads3mo agoHugging Face11sequelbox /Superpotion-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases! Superpotion-DeepSeek-V3.2.Speciale is a dataset containing structured medical reasoning responses, testing the limits of DeepSeek V3.2 Speciale's medical reasoning skills across a wide variety of medical disciplines and tasks! This dataset contains: 28.8k synthetically generated medical prompts, with all responses generated using DeepSeek V3.2 Speciale. Structured medical reasoning: Superpotion uses organized, informative… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Superpotion-DeepSeek-V3.2-Speciale.texttext-generation10K<n<100K7 likes79 downloads9mo agoHugging Face12sequelbox /DAG-Reasoning-DeepSeek-R1-0528Click here to support our open-source dataset and model releases! DAG-Reasoning-DeepSeek-R1-0528 is a dataset focused on analysis and reasoning, creating directed acyclic graphs testing the limits of DeepSeek R1 0528's graph-reasoning skills! This dataset contains: 4.08k synthetically generated prompts to create directed acyclic graphs in response to user input, with all responses generated using DeepSeek R1 0528. All responses contain a multi-step thinking process to perform effective… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DAG-Reasoning-DeepSeek-R1-0528.texttext-generation1K<n<10K12 likes76 downloads1y agoHugging Face13sequelbox /Titanium2.1-DeepSeek-R1Click here to support our open-source dataset and model releases! Titanium2.1-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills! This dataset contains: 31.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium2.1-DeepSeek-R1.texttext-generation10K<n<100K9 likes69 downloads1y agoHugging Face14sequelbox /Raiden-Mini-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases! Raiden-Mini-DeepSeek-V3.2.Speciale is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek-V3.2.Speciale's reasoning skills! This dataset contains: a default subset of ~8k 'creative_content' and 'analytical_reasoning' prompts from sequelbox/Raiden-DeepSeek-R1, with all responses generated by DeepSeek V3.2 Speciale. provides an unfiltered look into the reasoning skills of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-Mini-DeepSeek-V3.2-Speciale.texttext-generation10K<n<100K7 likes58 downloads10mo agoHugging Face15sequelbox /Mitakihara-DeepSeek-R1-0528Click here to support our open-source dataset and model releases! Mitakihara-DeepSeek-R1-0528 is a dataset focused on artificial intelligence, testing the limits of DeepSeek R1 0528's AI-reasoning skills! This dataset contains: 16.9k synthetically generated prompts about AI, with all responses generated using DeepSeek R1 0528. Subjects include computer science, artificial intelligence, MLOps, LLMs and diffusion models, math and CUDA, cutting-edge and future technologies, complex adaptive and… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Mitakihara-DeepSeek-R1-0528.texttext-generation10K<n<100K6 likes56 downloads1y agoHugging Face16sequelbox /UML-Generator-Dataset-DeepSeek-V3.2Click here to support our open-source dataset and model releases! UML-Generator-Dataset-DeepSeek-V3.2 is a dataset focused on analysis and code-reasoning, creating UML diagrams testing the limits of DeepSeek V3.2's modeling and design skills! This dataset contains: 2.7k synthetically generated prompts to create UML diagrams in response to user input, with all responses generated using DeepSeek V3.2. All responses contain a multi-step thinking process to perform effective analysis, followed by… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/UML-Generator-Dataset-DeepSeek-V3.2.texttext-generation1K<n<10K6 likes56 downloads11mo agoHugging Face17ansulev /deepseek-v4-pro-tachibana4Click here to support our open-source dataset and model releases - help us speed up our release schedule! Tachibana 4 is an agentic coding dataset, testing the limits of DeepSeek-V4-Pro's coding skills: Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Synthethic prompts utilize a variety of personas, experience levels, and styles of communication to maximize real-world flexibility and usability. Areas of focus include… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-pro-tachibana4.texttext-generation10K<n<100K0 likes54 downloads4mo agoHugging Face18gravermistakes /Titanium2-DeepSeek-R1Click here to support our open-source dataset and model releases! Titanium2-DeepSeek-R1 is a dataset focused on architecture and DevOps, testing the limits of DeepSeek R1's architect and coding skills! This dataset contains: 32.4k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek R1. Primary areas of expertise are architecture (problem solving, scenario analysis, coding, full SDLC) and DevOps (Azure, AWS, GCP, Terraform… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Titanium2-DeepSeek-R1.texttext-generation10K<n<100K0 likes53 downloads7mo agoHugging Face19sequelbox /Tachibana2-DeepSeek-R1Click here to support our open-source dataset and model releases! Tachibana2-DeepSeek-R1 is a code-reasoning dataset, testing the limits of DeepSeek R1's coding skills! This dataset contains: 27.2k synthetically generated code-reasoning prompts. All responses are generated using DeepSeek R1. Synthetic prompts are generated using Llama 3.1 405b Instruct, based on the original sequelbox/Tachibana dataset with increased task complexity. Responses demonstrate the code-reasoning capabilities of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana2-DeepSeek-R1.texttext-generation10K<n<100K5 likes51 downloads1y agoHugging Face20sequelbox /Tachibana4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule! This is an early sneak preview of Tachibana 4, containing the first 1.2k rows! Tachibana 4 is an upcoming agentic coding dataset, generated by DeepSeek-V4-Pro: Questions prioritize real-world, challenging agentic coding tasks across a variety of programming languages and topics. Areas of focus include back-end and front-end development, systems programming, distributed systems… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Tachibana4-DeepSeek-V4-Pro-PREVIEW.texttext-generation1K<n<10K17 likes51 downloads5mo agoHugging Face21MedSalim /QA-deepseek-r1-distill-llama-70b DeepSeek-R1-LLama-70B Q&A Dataset This repository contains a curated set of 484 questions and answers generated by the DeepSeek-R1-LLama-70B model. The main goal is to evaluate the quality, coherence, and factual correctness of the model’s responses under various scenarios. Before getting excited about it, let's be realistic—large language models can produce both impressive and abysmal results. This dataset is meant to help you figure out which side of that spectrum… See the full description on the dataset page: https://huggingface.co/datasets/MedSalim/QA-deepseek-r1-distill-llama-70b.textn<1K1 likes48 downloads2y agoHugging Face22sequelbox /Celestia3-DeepSeek-R1-0528-PREVIEWClick here to support our open-source dataset and model releases! This is an early sneak preview of Celestia3-DeepSeek-R1-0528, containing the first 13.4k rows! Celestia3-DeepSeek-R1-0528 is a dataset focused on science, testing the limits of DeepSeek R1's science-reasoning skills! This early preview release contains: 13.4k synthetically generated science prompts. All responses are generated using DeepSeek R1 0528. Primary subjects are physics, chemistry, biology, and computer science;… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Celestia3-DeepSeek-R1-0528-PREVIEW.texttext-generation10K<n<100K7 likes47 downloads1y agoHugging Face23mmrech /Superpotion-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases! Superpotion-DeepSeek-V3.2.Speciale is a dataset containing structured medical reasoning responses, testing the limits of DeepSeek V3.2 Speciale's medical reasoning skills across a wide variety of medical disciplines and tasks! This dataset contains: 28.8k synthetically generated medical prompts, with all responses generated using DeepSeek V3.2 Speciale. Structured medical reasoning: Superpotion uses organized, informative… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/Superpotion-DeepSeek-V3.2-Speciale.texttext-generation10K<n<100K1 likes47 downloads8mo agoHugging Face24gravermistakes /Titanium3-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases! Titanium3-DeepSeek-V3.1-Terminus is a dataset focused on architecture and DevOps, testing the limits of DeepSeek V3.1 Terminus's architect and coding skills! This dataset contains: 27.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek V3.1 Terminus in reasoning mode: 20k selected technical expertise prompts from sequelbox/Titanium2.1-DeepSeek-R1 focused on… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Titanium3-DeepSeek-V3.1-Terminus.texttext-generation10K<n<100K0 likes47 downloads7mo agoHugging Face25sequelbox /Raiden-DeepSeek-R1-PREVIEWThis is a preview of the full Raiden-Deepseek-R1 creative and analytical reasoning dataset, containing the first ~6k rows. Get the full dataset here! This dataset uses synthetic data generated by deepseek-ai/DeepSeek-R1. The initial release of Raiden uses 'creative_content' and 'analytical_reasoning' prompts from microsoft/orca-agentinstruct-1M-v1. Dataset has not been reviewed for format or accuracy. All responses are synthetic and provided without editing. Use as you will. text1K<n<10K6 likes42 downloads2y agoHugging Face26sequelbox /DES-Reasoning-DeepSeek-V3.1Click here to support our open-source dataset and model releases! DES-Reasoning-DeepSeek-V3.1 is a dataset focused on analysis and reasoning, creating discrete event simulations testing the limits of DeepSeek V3.1's simulation, Python scripting, and analysis skills! This dataset contains: 4.03k synthetically generated prompts to create discrete event simulations and analysis chat in response to user input, with all responses generated using DeepSeek V3.1. All responses contain a multi-step… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DES-Reasoning-DeepSeek-V3.1.texttext-generation1K<n<10K1 likes41 downloads1y agoHugging Face27ryen-stuff /Deepseek-code DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.tabulartext-generation1K<n<10K0 likes35 downloads1mo agoHugging Face28tuanha1305 /DeepSeek-R1-Distilltabular100K<n<1M11 likes34 downloads2y agoHugging Face29sequelbox /Titanium4-DeepSeek-V4-Pro-PREVIEWClick here to support our open-source dataset and model releases - help us speed up our release schedule! This is an early sneak preview of Titanium 4, containing the first 4.9k rows! Titanium 4 is an upcoming agentic coding dataset focused on DevOps and architecture, generated by DeepSeek-V4-Pro: Questions prioritize real-world, challenging agentic coding tasks in DevOps and architecture across a variety of programming languages and topics. Areas of focus include IaC, cloud architecture… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium4-DeepSeek-V4-Pro-PREVIEW.texttext-generation1K<n<10K1 likes34 downloads4mo agoHugging Face30lucsaint /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K0 likes34 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.