CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lvogel123 /jailbreak-deepseek-v3.2-exptabular1K<n<10K1 likes16k downloads11mo agoHugging Face02fireworks-ai /logiqa-deepseek-v3text1K<n<10K0 likes1k downloads2y agoHugging Face03OnDeviceMedNotes /synthetic-medical-conversations-deepseek-v3 🍎 Synthetic Multipersona Doctor Patient Conversations. Author: Nisten Tahiraj License: MIT 🧠 Generated by DeepSeek V3 running in full BF16. 🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival systems for reducing hallucinations and increasing the diagnosis quality. 🐧 Conversations… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/synthetic-medical-conversations-deepseek-v3.text35 likes938 downloads2y agoHugging Face04MaziyarPanahi /synthetic-medical-conversations-deepseek-v3-chatTaken from Synthetic Multipersona Doctor Patient Conversations. by Nisten Tahiraj. Original README 🍎 Synthetic Multipersona Doctor Patient Conversations. Author: Nisten Tahiraj License: MIT 🧠 Generated by DeepSeek V3 running in full BF16. 🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/synthetic-medical-conversations-deepseek-v3-chat.text1K<n<10K6 likes785 downloads2y agoHugging Face05NousResearch /eval-DeepSeek-V3-0324 dsv3 Evaluation Results Summary Benchmark Score Metric Samples Overlong rate aime24 0.506 math_pass@1:64_samples 64 100.0% aime25 0.422 math_pass@1:64_samples 64 100.0% arenahard 0.926 eval/overall_winrate 500 0.0% bbh_generative 0.868 extractive_match 1 100.0% creative-writing-v3 0.767 creative_writing_score 96 0.0% drop_generative_nous 0.829 drop_acc 1 100.0% eqbench3 0.831 eqbench_score 135 0.0% gpqa_diamond 0.680 gpqa_pass@1:8_samples8 100.0%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-V3-0324.tabular100K<n<1M2 likes219 downloads1y agoHugging Face06VINAY-UMRETHE /Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-Highgated Distill This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format. Dataset Structure The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.texttext-generation100K<n<1M14 likes86 downloads3mo agoHugging Face07Lego-X /Terminal-Lego-Traj-Deepseek-V3-2-15ktext10K<n<100K0 likes73 downloads4mo agoHugging Face08TeichAI /deepseek-v3.2-speciale-openr1-math-3kInspired by @OpenR1 The questions for this dataset were all sourced from the first 3.3k prompts in open-r1/OpenR1-Math-220k Dataset Stats (provided by OpenRouter): Cost: $ 21.1 (USD) Tokens (input + output): 52.3 M text1K<n<10K9 likes65 downloads10mo agoHugging Face09nebius /DeepSeek-V3-Infinity-Instruct-0625 DeepSeek-V3-Infinity-Instruct-0625 Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside DeepSeek-V3-0324 as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with deepseek-ai/DeepSeek-V3-0324 at temperature=1. For more details on the training methodology and… See the full description on the dataset page: https://huggingface.co/datasets/nebius/DeepSeek-V3-Infinity-Instruct-0625.texttext-generation100K<n<1M2 likes62 downloads7mo agoHugging Face10Aratako /Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k 概要 deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。 このデータセットはNSFW表現を含みます。 データの詳細 各データは以下のキーを含んでいます。 major_genre: ジャンル(大分類) minor_genre: ジャンル(小分類) tag: 年齢制限用タグ(R-18) world_setting: 舞台・世界観の設定 scene_setting: 対話シーンの設定 user_setting: ユーザー側のキャラクターの設定 assistant_setting: アシスタント側のキャラクターの設定 dialogue_tone: 対話のトーン conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式)… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k.texttext-generation10K<n<100K10 likes60 downloads1y agoHugging Face11wjn922-01 /scale-swe-distill5000-deepseek-v3.2-think-rollout4-instance1000-trajectories3368tabular1K<n<10K0 likes59 downloads2mo agoHugging Face12Aratako /Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k-formatted Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20k-formatted 概要 deepseek-ai/DeepSeek-V3-0324を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-NSFW-DeepSeek-V3-0324-20kにsystem messageを追加して整形したデータセットです。 データの詳細については元データセットのREADMEを参照してください。 ライセンス MITライセンスの元配布します。 text-generation10K<n<100K1 likes58 downloads1y agoHugging Face13sequelbox /UML-Generator-Dataset-DeepSeek-V3.2Click here to support our open-source dataset and model releases! UML-Generator-Dataset-DeepSeek-V3.2 is a dataset focused on analysis and code-reasoning, creating UML diagrams testing the limits of DeepSeek V3.2's modeling and design skills! This dataset contains: 2.7k synthetically generated prompts to create UML diagrams in response to user input, with all responses generated using DeepSeek V3.2. All responses contain a multi-step thinking process to perform effective analysis, followed by… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/UML-Generator-Dataset-DeepSeek-V3.2.texttext-generation1K<n<10K6 likes56 downloads10mo agoHugging Face14TeichAI /deepseek-v3.2-speciale-OpenCodeReasoning-3kThe questions for this dataset were all sourced from the first 3k prompts in nvidia/OpenCodeReasoning Dataset Stats (provided by OpenRouter): Cost: $ 19.2 (USD) Tokens (input + output): 47 M text1K<n<10K12 likes49 downloads10mo agoHugging Face15Jackrong /Chinese-DeepSeek-V3.2-Exp-chat-example deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本 一、前言 本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。 二、数据与方法 数据来源:用户构建的 6,655 轮真实中文对话样本。 估算方法: 中文字符近似为 1 Token; 英文 4 字符 ≈ 1 Token; 用于规模与上下文预算对比,而非精确 Token 计数。 统计维度: 平均 Prompt/Output 长度(字符与估算 Token); 总 Token 占上下文窗口比例; 语言分布(Prompt 语言类型); 对话长度分布(用户提问、助手回答、总对话长度)。 三、总体结果 1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.tabularquestion-answering1K<n<10K5 likes46 downloads1y agoHugging Face16Cetemadi1 /jailbreak-deepseek-v3.2-exptabular1K<n<10K0 likes46 downloads9mo agoHugging Face17sequelbox /Superpotion-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases! Superpotion-DeepSeek-V3.2.Speciale is a dataset containing structured medical reasoning responses, testing the limits of DeepSeek V3.2 Speciale's medical reasoning skills across a wide variety of medical disciplines and tasks! This dataset contains: 28.8k synthetically generated medical prompts, with all responses generated using DeepSeek V3.2 Speciale. Structured medical reasoning: Superpotion uses organized, informative… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Superpotion-DeepSeek-V3.2-Speciale.texttext-generation10K<n<100K7 likes46 downloads8mo agoHugging Face18mmrech /Superpotion-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases! Superpotion-DeepSeek-V3.2.Speciale is a dataset containing structured medical reasoning responses, testing the limits of DeepSeek V3.2 Speciale's medical reasoning skills across a wide variety of medical disciplines and tasks! This dataset contains: 28.8k synthetically generated medical prompts, with all responses generated using DeepSeek V3.2 Speciale. Structured medical reasoning: Superpotion uses organized, informative… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/Superpotion-DeepSeek-V3.2-Speciale.texttext-generation10K<n<100K1 likes40 downloads8mo agoHugging Face19gravermistakes /Titanium3-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases! Titanium3-DeepSeek-V3.1-Terminus is a dataset focused on architecture and DevOps, testing the limits of DeepSeek V3.1 Terminus's architect and coding skills! This dataset contains: 27.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek V3.1 Terminus in reasoning mode: 20k selected technical expertise prompts from sequelbox/Titanium2.1-DeepSeek-R1 focused on… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Titanium3-DeepSeek-V3.1-Terminus.texttext-generation10K<n<100K0 likes38 downloads7mo agoHugging Face20sequelbox /Raiden-Mini-DeepSeek-V3.2-SpecialeClick here to support our open-source dataset and model releases! Raiden-Mini-DeepSeek-V3.2.Speciale is a dataset containing creative-reasoning and analytic-reasoning responses, testing the limits of DeepSeek-V3.2.Speciale's reasoning skills! This dataset contains: a default subset of ~8k 'creative_content' and 'analytical_reasoning' prompts from sequelbox/Raiden-DeepSeek-R1, with all responses generated by DeepSeek V3.2 Speciale. provides an unfiltered look into the reasoning skills of… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Raiden-Mini-DeepSeek-V3.2-Speciale.texttext-generation10K<n<100K7 likes37 downloads10mo agoHugging Face21SWE-Router /swebench-verified-deepseek-v3.2tabularn<1K0 likes36 downloads5mo agoHugging Face22SWE-Router /v3-2k-traj-deepseek-v4-flashtabular1K<n<10K0 likes36 downloads5mo agoHugging Face23Aratako /Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k 概要 deepseek-ai/DeepSeek-V3-0324を用いて作成した、約20000件の日本語ロールプレイの対話を収録した合成データセットです。各データは10ターンから20ターン程度あります。 データの詳細 各データは以下のキーを含んでいます。 major_genre: ジャンル(大分類) minor_genre: ジャンル(小分類) tag: 年齢制限用タグ(全年齢、R-15) world_setting: 舞台・世界観の設定 scene_setting: 対話シーンの設定 user_setting: ユーザー側のキャラクターの設定 assistant_setting: アシスタント側のキャラクターの設定 dialogue_tone: 対話のトーン conversations: 上記設定に基づいたユーザーとアシスタントの対話(OpenAI messages形式) 設定等の情報からsystem… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k.texttext-generation10K<n<100K9 likes35 downloads1y agoHugging Face24Aratako /Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20k-formatted 概要 deepseek-ai/DeepSeek-V3-0324を用いて作成した日本語ロールプレイデータセットであるAratako/Synthetic-Japanese-Roleplay-SFW-DeepSeek-V3-0324-20kにsystem messageを追加して整形したデータセットです。 データの詳細については元データセットのREADMEを参照してください。 ライセンス MITライセンスの元配布します。 texttext-generation10K<n<100K2 likes35 downloads1y agoHugging Face25tonyshark /deepseek-v3-10kThe first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page for the full Pile for more info. Inspired by stas' great resource doing the same for OpenWebText textn<1K0 likes34 downloads2y agoHugging Face26dddraxxx /filtered_deepseek_v31_referring_expression_parsing DeepSeek v3.1 Quality-Filtered Referring Expression Parsing + Distractor Labels This dataset contains parsed referring expressions from the RefCOCO, RefCOCOg, and RefCOCO+ validation sets, processed using DeepSeek v3.1 with quality filtering, plus corresponding distractor label annotations in COCO format. Dataset Description Overview Model: DeepSeek v3.1 (deepseek-chat) Processing: Quality-filtered results from referring expression parsing Datasets: RefCOCO… See the full description on the dataset page: https://huggingface.co/datasets/dddraxxx/filtered_deepseek_v31_referring_expression_parsing.image-to-text10K<n<100K0 likes33 downloads1y agoHugging Face27simonycl /persuasiveness-leaderboard-inverted-deepseek_chat_v3224textn<1K0 likes32 downloads11mo agoHugging Face28dvilasuero /chemistry-reasoning-phi4-vs-deepseekv3textn<1K1 likes31 downloads2y agoHugging Face29sequelbox /DES-Reasoning-DeepSeek-V3.1Click here to support our open-source dataset and model releases! DES-Reasoning-DeepSeek-V3.1 is a dataset focused on analysis and reasoning, creating discrete event simulations testing the limits of DeepSeek V3.1's simulation, Python scripting, and analysis skills! This dataset contains: 4.03k synthetically generated prompts to create discrete event simulations and analysis chat in response to user input, with all responses generated using DeepSeek V3.1. All responses contain a multi-step… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/DES-Reasoning-DeepSeek-V3.1.texttext-generation1K<n<10K1 likes31 downloads1y agoHugging Face30sequelbox /Titanium3-DeepSeek-V3.1-TerminusClick here to support our open-source dataset and model releases! Titanium3-DeepSeek-V3.1-Terminus is a dataset focused on architecture and DevOps, testing the limits of DeepSeek V3.1 Terminus's architect and coding skills! This dataset contains: 27.7k synthetically generated prompts focused on architecture, cloud, and DevOps. All responses are generated using DeepSeek V3.1 Terminus in reasoning mode: 20k selected technical expertise prompts from sequelbox/Titanium2.1-DeepSeek-R1 focused on… See the full description on the dataset page: https://huggingface.co/datasets/sequelbox/Titanium3-DeepSeek-V3.1-Terminus.texttext-generation10K<n<100K2 likes31 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.