CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TeichAI /deepseek-v3.2-speciale-openr1-math-3kInspired by @OpenR1 The questions for this dataset were all sourced from the first 3.3k prompts in open-r1/OpenR1-Math-220k Dataset Stats (provided by OpenRouter): Cost: $ 21.1 (USD) Tokens (input + output): 52.3 M text1K<n<10K9 likes75 downloads10mo agoHugging Face02TeichAI /deepseek-v3.2-speciale-OpenCodeReasoning-3kThe questions for this dataset were all sourced from the first 3k prompts in nvidia/OpenCodeReasoning Dataset Stats (provided by OpenRouter): Cost: $ 19.2 (USD) Tokens (input + output): 47 M text1K<n<10K12 likes51 downloads10mo agoHugging Face03Jackrong /DeepSeek-v3.1-reasoner-Distilled-math-samples DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset) The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.tabularquestion-answeringn<1K1 likes33 downloads1y agoHugging Face04TeichAI /deepseek-v3.2-speciale-1000xThis is a reasoning dataset created using Deepseek v3.2 Speciale with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Deepseek v3.2 Speciale by fine-tuning already existing open-source LLMs. Stats Costs: $ 1.46 (USD) Total tokens (input + output): 3.61 M textn<1K16 likes33 downloads10mo agoHugging Face05Bouquets /DeepSeek-V3-Distill-Cybersecurity-en 🔐 CyberSecPentest-DS: Deepseek-V3 Distilled Dataset 📂 Dataset Overview:This is a high-quality distilled dataset 🧪, specialized in the cybersecurity penetration testing domain, generated using Deepseek-V3 🚀. It provides curated knowledge for red-teaming, vulnerability assessment, and ethical hacking research. 💡 Key Features: 🛡️ Focus: Penetration testing, exploit development, and security assessments. 🧠 Source: Knowledge distilled from Deepseek-V3 for accuracy &… See the full description on the dataset page: https://huggingface.co/datasets/Bouquets/DeepSeek-V3-Distill-Cybersecurity-en.text1K<n<10K0 likes27 downloads1y agoHugging Face06Jackrong /DeepSeek-V3.2-Exp-reasoning-example 🐳 DeepSeek-V3.2-Exp-reasoning vs DeepSeek-R1-0528: Math Reasoning Comparison 🍎 Note: DeepSeek-R1-0528 has no explicit chain-of-thought, while deepseek-ai/DeepSeek-V3.2-Exp (abbrev. V3.2-Exp) produces answers with structured derivations. This report was analyzed by GPT-5-Extended-Thinking. The sample size is small; conclusions are for reference only. Author: Soren 1. Executive Summary Sample size: 208 problems (mixed types). Average steps (reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V3.2-Exp-reasoning-example.tabularquestion-answeringn<1K2 likes26 downloads1y agoHugging Face07rAVEUK /deepseek-v3.2-speciale-OpenCodeReasoning-3kThe questions for this dataset were all sourced from the first 3k prompts in nvidia/OpenCodeReasoning Dataset Stats (provided by OpenRouter): Cost: $ 19.2 (USD) Tokens (input + output): 47 M text1K<n<10K0 likes22 downloads8mo agoHugging Face08EliotShen /Deepseek-V3-Distilled-Ancient-Chinese-Translation Dataset Card 这是一个文言文/白话文互译的高质量数据集,翻译精准,理由充分。总共有70K条数据,通过Deepseek-V3蒸馏获得。 Dataset Card Authors Shen ZhuoKang From ECNU Dataset Card Contact 10235101553@stu.ecnu.edu.cn texttranslation10K<n<100K3 likes18 downloads1y agoHugging Face09brandonbaek /DeepSeek-V3_Synthetic_Conversation_Dialogue DeepSeek V3 Synthetic Conversation Dialogue Dataset This dataset contains basic conversational dialogue for a chatbot with system prompts. Source DeepSeek V3 Model was used to generate this synthetic dataset. texttext-generationn<1K0 likes17 downloads2y agoHugging Face10Liontix /deepseek-v3.1-1000xThis is a reasoning dataset created using DeepSeek V3.1. Some of these questions are from reedmayhew and the rest were generated. Most of the questions cover the following topics: Web Development, Logic, Math, Embedded Systems, Web Design and Python Scripting. The dataset is meant for creating distilled versions of DeepSeek V3.1 by fine-tuning already existing open-source LLMs. textn<1K1 likes16 downloads1y agoHugging Face11Boakpe /deepseek-v3.2-thinking-html-distillation-750A small toy dataset generated using DeepSeek-V3.2-Thinking, designed for web design tasks involving HTML, CSS, and JavaScript. texttext-generationn<1K1 likes16 downloads8mo agoHugging Face12Liontix /deepseek-v3.1-200xThis is a reasoning dataset created using DeepSeek V3.1. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of DeepSeek V3.1 by fine-tuning already existing open-source LLMs. textn<1K0 likes12 downloads1y agoHugging Face13HashcakeML /claude-4.5-opus-high-reasoning-X-DeepseekV3.2-264xThis is a reasoning dataset created using Claude Opus 4.5 with a reasoning depth set to high. Some of these questions are from reedmayhew and the rest were generated. The dataset is meant for creating distilled versions of Claude Opus 4.5 by fine-tuning already existing open-source LLMs. Stats Costs: $ 52.3 (USD) Total tokens (input + output): 2.13 M textn<1K0 likes12 downloads9mo agoHugging Face14dianzinao /deepseek-v3-distill-RWA-1000 A Practical Exploration of Mixed-Style Response LLMs via Few-Shot LoRA Fine-Tuning 1. Research Background and Motivation With the increasingly widespread application of Large Language Models (LLMs) today, how to make model outputs more transparent, natural, and understandable has become an important research direction. Traditional LLMs typically output the final answer directly, making their internal reasoning process a "black box" to the user. To enhance the… See the full description on the dataset page: https://huggingface.co/datasets/dianzinao/deepseek-v3-distill-RWA-1000.text1K<n<10K0 likes9 downloads11mo agoHugging Face15alsim-01 /TE_dataset_deepseekv3_2.jsonltextn<1K0 likes7 downloads6mo agoHugging Face16NewEden-Forge /Deepseek-V3-RPtext1K<n<10K0 likes5 downloads1y agoHugging Face17dianzinao /deepseek-v3-distill-RWA-10000text10K<n<100K0 likes5 downloads11mo agoHugging Face18Fizzarolli /deepseek-v3neo-synthkinkgatedgweep gwoop text1K<n<10K0 likes2 downloads1y agoHugging Face19jasong03 /health_deepseekV3gatedtextn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.