datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Medical-R1-Distill-Data
Introduction
This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1.
The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese.
The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.Chinese-DeepSeek-R1-Distill-data-110k
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。
该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.Chinese-DeepSeek-R1-Distill-data-110k-SFT
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下:
Math:共计36568个样本,
Exam:共计2432个样本,
STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.Medical-R1-Distill-Data-Chinese
Introduction
This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on Chinese medical verifiable problems from HuatuoGPT-o1.
The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long reasoning chains based on GPT-4o on medical-o1-reasoning-SFT.
For details, see our paper and GitHub… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data-Chinese.Chinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.bielik-distill-polish-10k
bielik-distill-polish-10k
Polish instruction-tuning dataset with 10,304 samples generated via response-level knowledge distillation from Bielik-11B-v3.0-Instruct (SpeakLeash, Apache 2.0).
Covers Polish history, culture, politics, science, geography, idioms, and general reasoning. Multi-pass quality control: factual corrections, topic filtering (Poland/Europe focus), truncation removal (~9% of raw data removed).
Format
{
"messages": [
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/JohnTdi/bielik-distill-polish-10k.gpt-oss-120B-distilled-reasoning
GPT-oss-120B-Distilled-Reasoning-math Dataset
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines
Fields: Generator, Category, Input, Output
Core Statistics
Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and Answer.To understand the data… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-reasoning.GPT-OSS-120B-Distilled-Reasoning-math
GPT-oss-120B-Distilled-Reasoning-math Dataset
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines
Fields: Generator, Category, Input, CoT_Native_Reasoning, Reasoning, Answer
Core Statistics
Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-120B-Distilled-Reasoning-math.MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B.
The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning.
We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details.
If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.hwtcm-deepseek-r1-distill-data
简介
DeepSeek蒸馏的传统中医数据集,原始数据来源于网络,未进行人工审查。
7B模型微调效果
模型表现出了推理能力,准确性有待继续验证。
我们的其他产品
中医NER:能识别方剂、本草、来源、病名、症状、证型,也许是基于BERT开源模型中识别最好的模型。中医考试题:也许是全网最早开源、数据最多的中医考试题,我们内部将其用于模型训练的性能评测数据集。中医SFT数据集:中医QA数据集,用于SFT微调。仓公:基于Qwen的指令微调模型(暂未开源)。仓公R1:基于DeepSeek蒸馏的超过100万条QA的指令微调模型,拥有强大的推理能力(暂未开源)。
。。。还有很多
Citation
If you find this project useful in your research, please consider cite:
@misc{hwtcm2024,
title={{hwtcm-deepseek-r1-distill-data} A traditional… See the full description on the dataset page: https://huggingface.co/datasets/Monor/hwtcm-deepseek-r1-distill-data.DeepSeek-R1-Distill-Qwen-32B-LeaPPaper: Learning from Peers in Reasoning Models
Project Page: https://learning-from-peers.github.io/
Code: https://github.com/tongxuluo/LeaP
GPT-OSS-20B-Distilled-Reasoning-Mini
Dataset Card for Dataset Name
GPT-OSS-20B Distilled Reasoning Dataset Mini
(Multi-stage Evaluative Refinement Method for Reasoning Generation)
Dataset Details and Description
This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.Qwen3-235B-A22B-Instruct-2507-Distilled-chat
Qwen3-235B-A22B-Instruct-2507-Distilled-chat📚
Curated/Funded/Shared by: [Jack Rong]
Language(s): English (major), Chinese, Русский, 한국어, 日本語, others
License: [apache-2.0]
Distilled Model: 🏆Qwen/Qwen3-235B-A22B-Instruct-2507
Qwen3-235B-A22B-Instruct-2507 Benchmarks📊
Introduction:
The objectives of this project are:
Focus on chat capabilities (excluding CoT), covering cross-lingual real-world Q&A/explanation/generation;
Utilize… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Qwen3-235B-A22B-Instruct-2507-Distilled-chat.Thai-R1-Distill-SFT
Thai R1 Distill SFT
Thai Reasoning Dataset for Supervised Finetuning
Translated by iApp Technology
gpt-oss-120B-distilled-math-OpenAI-Harmony
📚 Dataset Overview
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines (.jsonl)Fields: Generator, Category, Input, Output
Note: If you are using this template for training, please make sure the format is correct before starting.Since this template is still under continuous improvement and learning, it may not be fully complete yet. I appreciate your understanding.
📈 Core Statistics
Generated complete reasoning processes… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-math-OpenAI-Harmony.medqa-distill-sft
MedQA-Distill-SFT
医学多选题 → 中文推理 SFT 数据(由 MedQA 题库经 LLM API 蒸馏生成)
从 MedQA(美国 USMLE / 中国执业医师考试)多选题蒸馏而来:每题包含题目、选项、逐步临床推理(中文) 和答案。用于微调中文医学大模型(SFT / LoRA)。
数据统计
字段
数值
总条数
39,618
训练集
35,656
验证集
3,962
来源 Provider
DeepSeek / TokenRhythm / Agnes-CN / NVIDIA
推理语言
中文(英文题也生成中文推理)
数据格式(Alpaca)
{
"instruction": "题目:A 23-year-old pregnant woman...\n选项:\nA. ...\nB. ...\nC. ...\nD. ...\nE. ...",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/rewrewrv343/medqa-distill-sft.AtmosphericQA-1k-Chinese-Distillation
AtmosphericQA-distillation
Overview
AtmosphericQA-distillation is a Chinese supervised fine-tuning (SFT) question–answering dataset focused on atmospheric science and meteorology.The dataset is constructed via knowledge distillation from the Gemini 3 Flash Preview model, with the goal of providing systematic, structured, and domain-specific scientific knowledge for Chinese large language models.
It covers a broad range of subfields, from fundamental atmospheric theory to… See the full description on the dataset page: https://huggingface.co/datasets/phoenixcph/AtmosphericQA-1k-Chinese-Distillation.pocket-mechanic-distilled
Pocket Mechanic: distillation dataset
4,365 (sensor window → diagnostic explanation) pairs for fine-tuning small language models to do OBD-II car-fault diagnosis with calibrated repair-cost estimates and shop-upsell awareness.
Used to train MindFreakGamer/gemma-4-E2B-pocket-mechanic. The student reaches 81.3% of teacher quality on a blind A/B judge (n=100).
Submission to the Hugging Face Build Small Hackathon, Backyard AI track (June 2026). Code:… See the full description on the dataset page: https://huggingface.co/datasets/MindFreakGamer/pocket-mechanic-distilled.Qwen3-Reasoning-Distill-Q-A-Dataset
Qwen3 Reasoning Distill Q&A Dataset
Repository: RefinedNeuro/Qwen3-Reasoning-Distill-Q-A-Dataset
Authors
Mehmet Can Farsak
Serhat Atayeter
License
This dataset is released under CC0 1.0 Universal (CC0 1.0) Public Domain Dedication.
Dataset Summary
This dataset contains question-answer pairs across six STEM subjects designed for Turkish-language reasoning tasks. It was generated using the qwen3-32b model and is intended for fine-tuning the RN_TR_R2… See the full description on the dataset page: https://huggingface.co/datasets/RefinedNeuro/Qwen3-Reasoning-Distill-Q-A-Dataset.qwen3-coder-480b-distill-mini
qwen3-coder-480b-distill-mini
Short Description
This dataset is distilled using Qwen3-Coder-480B-A35B-Instruct.We extracted 10,000 code questions from microsoft/rStar-Coder as seed problems, distilled them with 32K context, and after cleaning and filtering, 9,543 samples remain.License: Apache-2.0.
Dataset Overview
Seed Source: 10,000 code reasoning problems sampled from microsoft/rStar-Coder.
Distillation Model: Qwen3-Coder-480B-A35B-Instruct (480B… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/qwen3-coder-480b-distill-mini.medical_r1_distill_sft_Chinese_alpacaBorrowed from https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data-Chinese
Reorganized the data structure for easier SFT.
DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.faithful-tom-distillation
Dataset for "Faithful Theory of Mind Distillation"
This repository contains the datasets used in the paper "Faithful Theory of Mind Distillation: Why Preference Based Refinement Improves Imitation", accepted to the AAAI 2026 ToM4AI Workshop.
Dataset Structure
The repository contains two files corresponding to the two training stages described in the paper:
train_sft_combined.jsonl:
Purpose: Used for the Supervised Fine-Tuning (SFT) stage.
Content: Contains social… See the full description on the dataset page: https://huggingface.co/datasets/ArpitSinghGautam/faithful-tom-distillation.Deepseek-V3-Distilled-Ancient-Chinese-Translation
Dataset Card
这是一个文言文/白话文互译的高质量数据集,翻译精准,理由充分。总共有70K条数据,通过Deepseek-V3蒸馏获得。
Dataset Card Authors
Shen ZhuoKang From ECNU
Dataset Card Contact
10235101553@stu.ecnu.edu.cn
gpt2_to_gpt5.5_distilled_25k
GPT-2 to GPT-5.5 Advanced Reasoning Distillation (25k)
Dataset Description
25,000 unique, high-quality instruction-response pairs designed for knowledge distillation and supervised fine-tuning. The dataset elevates GPT-2 Medium toward GPT-5.5-level performance on complex reasoning tasks.
Core goal: Transfer frontier reasoning capabilities (multi-step CoT, cross-domain synthesis, edge-case analysis, novel insights) from a hypothetical GPT-5.5 teacher into smaller… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gpt2_to_gpt5.5_distilled_25k.Chinese-DeepSeek-R1-Distill-data-110k
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。
该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/Chinese-DeepSeek-R1-Distill-data-110k.distill_r1_110k_sft_modifiedBorrowed from https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT
Fix the <image> placeholder issue, which will cause error during training:
raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.")
crimeopus-distill-v2
CrimeOpus 4.7 — Distilled Coding Dataset
Dataset di fine-tuning per CrimeOpus 4.7-v2 (LoRA training).
Sources
Source
Count
Type
DeepSeek-Chat distillation
157
Multi-domain coding/reasoning
Git commit-diff (CrimeCode-IDE)
91
Real codebase patterns
Uncensored seed (toxic-dpo + orpo-mix)
179
Refusal-free helpfulness
Total
427
Format
ChatML messages array:
{
"messages": [
{"role": "system", "content": "..."},
{"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/JollyFraud/crimeopus-distill-v2.DEEPMIND_Alpha_Distilled
AlphaMirror-25k: DeepMind Alpha-Inspired Reasoning Dataset for LLM Fine-Tuning
Dataset Summary
AlphaMirror-25k is a high-quality synthetic dataset containing exactly 25,000 instruction-response pairs designed to fine-tune any large language model to mirror the advanced reasoning, scientific discovery, and problem-solving style of Google DeepMind's Alpha series (AlphaFold, AlphaGo/AlphaZero, AlphaEvolve, AlphaGeometry, AlphaTensor, and related systems).
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DEEPMIND_Alpha_Distilled.DeepSeek-R1-Distill-Data-5k
