CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentica-org /DeepScaleR-Preview-Dataset Data Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from: AIME (American Invitational Mathematics Examination) problems (1984-2023) AMC (American Mathematics Competition) problems (prior to 2023) Omni-MATH dataset Still dataset Format Each row in the JSON dataset contains: problem: The mathematical question text, formatted with LaTeX notation. solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.text10K<n<100K206 likes39k downloads2y agoHugging Face02efficient-deep-research /synthesized_datasettext10K<n<100K0 likes34k downloads11mo agoHugging Face03deepguess /wrf-250m-severe-weather WRF 250m Severe Weather Simulations — Raw WRF Output 2 TB of raw WRF-ARW v4.7.1 output from 18 high-impact weather events at 250m resolution. Three nested domains (3km / 1km / 250m), 81 vertical levels, 216 variables per wrfout file, 1-minute surface output via auxhist2. This is the full simulation output — nothing removed, nothing derived. If you want pre-extracted surface fields in a simpler format, see wrf-250m-severe-weather-netcdf4 (16 fields, ~150 GB). This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/deepguess/wrf-250m-severe-weather.textothern<1K0 likes16k downloads7mo agoHugging Face04TeichAI /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K107 likes6.1k downloads4mo agoHugging Face05MoreThought /DeepSWEGym2 Dataset Description This dataset is a filtered and deduplicated version of a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 214.19kb, a total uncompressed size of 17.56GB, and a total of 85974 examples. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2.texttext-generation100K<n<1M1 likes2.3k downloads16d agoHugging Face06a-m-team /AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training. Due to certain constraints, we are only able to open-source a subset of the complete dataset. Model Training Performance based on our complete dataset On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.tabulartext-generation10M<n<100M56 likes2.1k downloads1y agoHugging Face07MoreThought /DeepSWEGym2-Ultra Dataset Description This dataset is an EXTREMELY filtered and deduplicated version of DeepSWE-Gym2-Edu, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 8.27MB, a total uncompressed size of 8.27GB, and a total of 1000 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Ultra.texttext-generationn<1K3 likes1.4k downloads16d agoHugging Face08ronaldcmz /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K0 likes1.1k downloads3mo agoHugging Face09Rock23210 /AIME_Deepseek_Cleantextn<1K0 likes1.1k downloads2y agoHugging Face10Hugodonotexit /math-code-science-deepseek-r1-en R1 Dataset Collection Aggregated high-quality English prompts and model-generated responses from DeepSeek R1 and DeepSeek R1-0528. Dataset Summary The R1 Dataset Collection combines multiple public DeepSeek-generated instruction-response corpora into a single, cleaned, English-only JSONL file. Each example consists of a <|user|> prompt and a <|assistant|> response in one "text" field. This release includes: ~21,000 examples from the DeepSeek-R1-0528 Distilled Custom… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/math-code-science-deepseek-r1-en.textquestion-answering1M<n<10M5 likes1.1k downloads1y agoHugging Face11Jackrong /DeepSeek-V4-Pro-Distilled-200K DeepSeek‑V4‑Pro‑Distilled‑200K High-quality Math & STEM reasoning distilled from DeepSeek‑V4‑Pro in Max mode Reasoning traces · Proofs · Verification · Mathematics · Physics · Chemistry · Biology Overview DeepSeek‑V4‑Pro‑Distilled‑200K is a supervised fine-tuning collection of long-form mathematical and scientific reasoning. Its responses were generated with DeepSeek‑V4‑Pro in Max inference mode, then normalized into a compact conversational… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Pro-Distilled-200K.texttext-generation100K<n<1M10 likes991 downloads2mo agoHugging Face12MoreThought /DeepSWEGym2-Edu Dataset Description This dataset is a filtered and deduplicated version of a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with complex/long code problems in the original datasets, having an average row size of 291.74kb, a total uncompressed size of 14.93GB, and a total of 53649 examples. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Edu.texttext-generation10K<n<100K2 likes988 downloads17d agoHugging Face13deepseek-ai /DeepSeek-Prover-V1 Evaluation Results | Model & Dataset Downloads | License | Contact Paper Link👁️ DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 1. Introduction Proof assistants like Lean have revolutionized mathematical proof verification, ensuring high accuracy and reliability. Although large language models (LLMs) show promise in… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-Prover-V1.text10K<n<100K74 likes899 downloads2y agoHugging Face14Congliu /Chinese-DeepSeek-R1-Distill-data-110k 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face&nbsp;&nbsp; | &nbsp;&nbsp;🤖 ModelScope &nbsp;&nbsp; | &nbsp;&nbsp;🚀 Github &nbsp;&nbsp; | &nbsp;&nbsp;📑 Blog 注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。 该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.tabulartext-generation100K<n<1M789 likes839 downloads2y agoHugging Face15MoreThought /DeepSWEGym2-Full Dataset Description This dataset is a merge containing many high quality SWE datasets, it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 169.08kb, a total uncompressed size of 19.80GB, and a total of 122791 examples. Dataset Details Curated by: MoreThought Funded by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym2-Full.texttext-generation100K<n<1M1 likes829 downloads16d agoHugging Face16Jackrong /DeepSeek-V4-Distill-8000x 🐳 DeepSeek-V4-Distill-8100x Dataset Summary DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash. After the cleaning process, the released train split contains 7,716 high-quality JSONL examples. [!NOTE] The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V4-Distill-8000x.texttext-generation1K<n<10K93 likes744 downloads5mo agoHugging Face17MoreThought /DeepSWEGym-Edu Dataset Description This dataset is a heavily filtered version of all the SWE-bench/SWE-smith-lang datasets (expect php) merged together (originally 88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is specifically filtered for rows with very complex/long code problems in the original datasets, having an average row size of 101.81kb, a total uncompressed size of 4.29GB, and 42529 examples total.… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Edu.texttext-generation10K<n<100K2 likes693 downloads17d agoHugging Face18MoreThought /DeepSWEGym-Full Dataset Description This dataset is a merged version of all the SWE-bench/SWE-smith-lang datasets (88k rows total), it aims to improve benchmark results on DeepSWE-style problems, benchmarks, and general coding skills. It is NOT specifically filtered for rows with complex/long code problems in the original datasets, despite still having an average row size of 85.4kb, a total uncompressed size of 7.53GB, and 88130 examples total. Dataset Details Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/MoreThought/DeepSWEGym-Full.texttext-generation10K<n<100K1 likes680 downloads17d agoHugging Face19hardcoremoore /DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. DeepSeek v4 Pro Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by deepseek/deepseek-v4-pro. JSONL files: 4006 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/hardcoremoore/DeepSeek-v4-Pro-Agent.tabulartext-generation1K<n<10K2 likes645 downloads4mo agoHugging Face20deepseek-ai /DeepSeek-ProverBench 1. Introduction We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with initialization data collected through a recursive theorem proving pipeline powered by DeepSeek-V3. The cold-start training procedure begins by prompting DeepSeek-V3 to decompose complex problems into a series of subgoals. The proofs of resolved subgoals… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-ProverBench.textn<1K47 likes609 downloads1y agoHugging Face21Deepexi /glaive-function-calling-vicuna数据集格式说明: glaiveai/glaive-function-calling · Datasets at Hugging Face 的 SFT 格式 我们高兴地宣布,数据集 "glaiveai/glaive-function-calling" 已经根据 SFT(Supervised Fine-Tuning)的需求进行了格式转换,以支持大型语言模型的训练。以下是有关这一新格式的简要说明: 数据集概述: 数据集 "glaiveai/glaive-function-calling" 基于 CC-BY-4.0 协议发布,原始数据集包含标识符和对话信息,这些数据已被转换为适应 SFT 训练的结构。 数据格式: 转换后的数据集格式包含以下关键信息: id: 整数类型的标识符,用于唯一标识每个数据样本。 conversations: 一个数组,其中包含对话信息。每个对话可以由多个句子组成,以更好地呈现函数调用的上下文。 数据集用途:转换后的数据集适用于 SFT 的训练,主要用途包括但不限于: 函数调用理解:… See the full description on the dataset page: https://huggingface.co/datasets/Deepexi/glaive-function-calling-vicuna.text10K<n<100K5 likes595 downloads3y agoHugging Face22ScienceOne-AI /S1-DeepResearch-15k S1-DeepResearch-15k Dataset Overview The S1-DeepResearch dataset is a curated collection of approximately 15k samples designed to improve deep research capabilities of large language models. The dataset includes two types of tasks: Verifiable tasks (labeled as "Closed-ended Multi-hop Resolution") Open-ended tasks (labeled as "Open-ended Exploration") Dataset Composition The dataset is organized into five core capability dimensions: Long-chain complex… See the full description on the dataset page: https://huggingface.co/datasets/ScienceOne-AI/S1-DeepResearch-15k.text10K<n<100K12 likes590 downloads5mo agoHugging Face23LocalLLaMA /deepswe-mini deepswe-mini 16 of the 113 tasks in DeepSWE v1.1, picked so that running just these ranks models the same way the full benchmark does. DeepSWE is a good benchmark and an expensive one. Every task is a long-horizon feature request in its own container, and a full pass takes close to two days of agent time run one task at a time. If you are comparing models, agent harnesses or prompts, and the differences you care about are more than a few points, these 16 tasks give you the same… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/deepswe-mini.tabularn<1K5 likes574 downloads7d agoHugging Face24SWE-Factory /DeepSWE-Agent-Kimi-K2-Trajectories-2.8Ktext1K<n<10K8 likes553 downloads1y agoHugging Face25danikhan632 /DeepMath-12ktext1K<n<10K0 likes515 downloads9mo agoHugging Face26DeepNLP /NIPS-2020-Accepted-Papers Neural Information Processing Systems NeurIPS 2020 Accepted Paper Meta Info Dataset This dataset is collected from the NeurIPS 2020 Advances in Neural Information Processing Systems 35 conference accepted paper (https://papers.nips.cc/paper_files/paper/2020) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/nips2020). For researchers who are interested in doing analysis of NIPS 2020 accepted papers and potential research trends, you can use the… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/NIPS-2020-Accepted-Papers.texttext-classification1K<n<10K0 likes504 downloads2y agoHugging Face27deep-principle /science_chemistrytextn<1K2 likes445 downloads15h agoHugging Face28Congliu /Chinese-DeepSeek-R1-Distill-data-110k-SFT 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face   |   🤖 ModelScope    |   🚀 Github    |   📑 Blog 注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下: Math:共计36568个样本, Exam:共计2432个样本, STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.tabulartext-generation100K<n<1M225 likes439 downloads2y agoHugging Face29AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes432 downloads6d agoHugging Face30deepguess /wrf-250m-severe-weather-netcdf4 WRF 250m Severe Weather Simulations — CF-Compliant NetCDF4 18 high-resolution (250m) WRF-ARW severe weather simulations in standard CF-compliant NetCDF4 format. Ready for immediate use with xarray, MetPy, Panoply, QGIS, GrADS, NCL, CDO, and any CF-aware tool. This is the NetCDF4 version of wrf-250m-severe-weather-overlays. Same data, standard format. Quick Start import xarray as xr ds = xr.open_dataset("carr_fire.nc") print(ds) # Plot composite reflectivity at… See the full description on the dataset page: https://huggingface.co/datasets/deepguess/wrf-250m-severe-weather-netcdf4.textothern<1K0 likes397 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.