datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KodCode-V1-SFT-R1
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.KodCode-V1-SFT-4o
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning.
🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-4o.bhasha-sft
Bhasha SFT
Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual
Large Language Models. The dataset contains collation of over 13 million instances of
instruction-response data for 3 Indian languages (Hindi, Gujarati, Bengali) and English having both human annotated and synthetic data.
Curated by: Soket AI Labs
Language(s) (NLP): [English, Hindi, Bengali, Gujarati]
License: [cc-by-4.0, apache-2.0, mit]… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/bhasha-sft.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.Chinese-DeepSeek-R1-Distill-data-110k-SFT
中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1)
🤗 Hugging Face | 🤖 ModelScope | 🚀 Github | 📑 Blog
注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。
本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。
为什么开源这个数据?
R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。
为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下:
Math:共计36568个样本,
Exam:共计2432个样本,
STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.turkce-sft-qa-3.7m
🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti
3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde
tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri
setinden geldiğini taşır.
English: A merged, row-level deduplicated and quality-filtered collection of
24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries
its source dataset, source URL and original license.
🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.synthetic-pre1930-sftTL;DR
A vintage finetuning dataset (~416k rows, eleven task routes). Sourced by taking excerpts
from pre-1930's texts, turning these into verbatim answers, and then using deepseek-chat to
generate period-appropriate questions of those answers. Any model tuned on this dataset should,
theoretically, never update its weights on anachronistic text, since questions are masked in the
finetuning stages. Features composition, verse, narrative, reasoning, multiturn dialogue, and
calibrated uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/synthetic-pre1930-sft.korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.FineProofs-SFT
FineProofs SFT
Dataset Description
FineProofs SFT is a high-quality supervised fine-tuning dataset containing mathematical Olympiad problems paired with chain-of-thought reasoning and formal proofs distilled from DeepSeek-Math-V2. The dataset comprises 7,777 samples (4,300 unique problems) sourced from international Olympiad competitions and Art of Problem Solving (AoPS), each annotated with:
Detailed reasoning traces (thinking content) generated by… See the full description on the dataset page: https://huggingface.co/datasets/lm-provers/FineProofs-SFT.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.chinese-legal-sft
Chinese Legal SFT Dataset(中文法律 SFT 数据集)
面向大模型监督微调(SFT)的中文法律问答数据集,共 19,332 条问答对,
每条附带 LLM 质量评分。覆盖数据采集 → 清洗 → 去重 → 质量过滤 → 格式化 → 质量打分的完整数据工程流程。
配套代码与完整流水线:https://github.com/noah-white-python/legal-sft-dataset
数据构建流程
冷启动:基于开源数据集 DISC-Law-SFT 整理。
清洗:NFKC 全角半角统一、去控制字符、去空白、缺失过滤。
去重:精确去重(MD5)+ MinHash + LSH 近似去重(阈值 0.8)。
质量过滤:长度、中文字符占比等启发式规则,有效率 96.7%(20,000 → 19,332)。
格式化:输出标准 Alpaca 指令格式。
质量打分:用 LLM-as-judge 对全部数据从复杂度、清晰度、信息量三维度打分(1-5 分)。
字段说明… See the full description on the dataset page: https://huggingface.co/datasets/noah248/chinese-legal-sft.toy-models-of-sft-data
Toy Models of SFT Data
This is a public-clean candidate data package for the Toy Models of SFT project.
It is built for researcher inspection first.
The package answers two questions:
What were the models trained on?
How did the models actually behave under evaluation?
The package includes training data, eval inputs, model rollouts, judge scores,
parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and
provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.gigaverbo-v2-rec-sft
GigaVerbo-v2 REC SFT
A model should not merely know how to reason; it should learn when reasoning is worth the cost.
Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B
Dataset Summary
GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.Maux-Persian-SFT-30k
Maux-Persian-SFT-30k
Dataset Description
This dataset contains 30,000 high-quality Persian (Farsi) conversations for supervised fine-tuning (SFT) of conversational AI models. The dataset combines multiple sources to provide diverse, natural Persian conversations covering various topics and interaction patterns.
Dataset Structure
Each entry contains:
messages: List of conversation messages with role (user/assistant/system) and content
source: Source dataset… See the full description on the dataset page: https://huggingface.co/datasets/xmanii/Maux-Persian-SFT-30k.tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.LogosForge-scored-sft-v1
LogosForge-scored-sft-v1
This dataset is a scored Supervised Fine-Tuning (SFT) distillation corpus built on top of the Natural Reasoning question set.It is constructed in two stages:
First, a large-scale reasoning-oriented teacher model (gpt-oss-120B-high) is used to generate distilled student responses, including explicit chain-of-thought reasoning, for natural reasoning questions.
Second, these distilled responses are evaluated by a separate instruction-following model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogosForge-scored-sft-v1.k3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.Vietverse-SFT-1K-Gold
🇻🇳 Vietverse-SFT (1K Gold Edition)
The "Less is More" Alignment Paradigm for Native Vietnamese Large Language Models
Bộ Dữ Liệu SFT Tiếng Việt Bản Xứ 1.000 Mẫu Gold Tinh Hoa — Chuẩn Mực Căn Chỉnh Mô Hình Ngôn Ngữ
🇻🇳 [Đọc Báo Cáo Kỹ Thuật Tiếng Việt] •
🇬🇧 [Read English Technical Card]
🤗 Hugging Face Dataset • ⚡ Hướng Dẫn Huấn Luyện / Quickstart
🌐 Ngôn Ngữ / Language
📌 Chuyển Hướng Nhanh / Quick Jump… See the full description on the dataset page: https://huggingface.co/datasets/TTP01/Vietverse-SFT-1K-Gold.MLR_sft_data
MLR SFT Data
MLR SFT Data is a teacher-generated supervised fine-tuning dataset for training Multi-Level Reasoning (MLR) models in the paper Enhancing Language Model Reasoning with Structured Multi-Level Modeling (ICLR 2026). It decomposes complete reasoning trajectories into two types of step-level examples:
Planner: plans the next reasoning goal and task based on the problem and reasoning history.
Executor: executes the Planner's instruction and updates the reasoning state.… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/MLR_sft_data.SimScholar-SFT
S3 SFT Trajectories
Complete ReAct trajectories for scientific-literature search.
Code ·
S3 collection ·
Source corpus
This dataset contains 14,633 single- and two-hop tool-use trajectories. In each
trajectory, a policy searches and reads a fixed scientific corpus through nine
tools, then submits an answer with a correctness label. The messages column
uses OpenAI tool-calling chat format.
At a glance
Question type
Rows
Correct
Incorrect
Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.LogicMind-Chat-Reasoning-SFT-300K
Nemotron-Post-Training-Dataset-v2-chat Dataset Card
Overview 📌
This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line).
Highlights
Scale: 296,168 samples
Category: chat (100%)
Generator: qwen-3-32b (100%)
Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.opc-sft-stage2-dense-extracted
OpenCoder Dataset Dense Region Extracted
This dataset is a post-processed version of the OpenCoder SFT Stage2 dataset (opc-sft-stage2).
We use gpt-4o API to extract the information dense regions from each sample and logged them in the dense_snippets column.Detailed information about the data can be found in our paper.
OpenCoder's sft-stage2 summary
The original version of this dataset is used in OpenCoder's Stage 2 and consists of four parts:
educational_instruct:… See the full description on the dataset page: https://huggingface.co/datasets/malr07/opc-sft-stage2-dense-extracted.nexttoken-pmkisan-domain-sft-data
NextToken pmkisan domain SFT data (v1)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
~57 scheme/product source documents (PM-KISAN, Ayushman Bharat, MGNREGA,
banking products, insurance, savings instruments, etc.).
Files… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-pmkisan-domain-sft-data.sft-data-1
sft-data-1
This dataset contains 100 SFT trajectories for DABench-style data-agent training.
Contents
train_sft.jsonl: JSONL SFT records.
train_sft.parquet: Parquet version of the same records.
summary.json: generation and filtering summary.
Generation
Expert model metadata: gpt-5.5
Tool/action format: qwen3_xml
Model judge: glm-5.2
Model-judge completed: 266
Model-judge accepted: 143
Selected SFT rows: 100
Assistant actions are serialized as… See the full description on the dataset page: https://huggingface.co/datasets/forseasons/sft-data-1.Vietverse-SFT-Preview
🇻🇳 Vietverse-SFT v1.0 (Preview Edition)
A High-Density, Native Vietnamese Foundation Dataset for Supervised Fine-Tuning
Bộ Dữ Liệu Nền Tảng SFT Tiếng Việt Bản Xứ Chuẩn Mực Cho Huấn Luyện Mô Hình Ngôn Ngữ Lớn
🇻🇳 [Đọc Bản Tiếng Việt] •
🇬🇧 [Read English Version]
🤗 Hugging Face Dataset • ⚡ Hướng Dẫn Sử Dụng / Quickstart
🌐 Ngôn Ngữ / Language
📌 Chuyển Hướng / Quick Navigation
🇻🇳 Bản Tiếng Việt Đầy Đủ… See the full description on the dataset page: https://huggingface.co/datasets/TTP01/Vietverse-SFT-Preview.faqih_sft_dataset
💎 MAFQA: Perfected Multi-Hop Arabic Fatwa QA Dataset (388 Samples)
مجموعة بيانات الاستدلال الفقهي المركب ومتعدد الخطوات (جامعة الملك سعود / MDPI 2026)
100% Curated & Unabridged MAFQA Multi-Hop Dataset
تم تنقيح وتدقيق البيانات بالكامل:
1. إزالة جميع التقطيعات النصية وإيراد النصوص والأدلة كاملة دون بتر.
2. تفعيل خطوة التركيب والترجيح النهائي (الخطوة 4) بربط استدلالي حقيقي بين المسائل الفرعية.
3. تصحيح الأخطاء المطبعية في دلالات الحل والحرمة.… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/faqih_sft_dataset.ipw-sft-trajectories
IPW SFT Trajectories
Supervised fine-tuning trajectories for training tool-use orchestrator models. Each trajectory contains a multi-turn conversation where a model solves a task by selecting and using tools (calculator, code interpreter, think).
Dataset Details
Stat
Value
Total trajectories (deduped, correct only)
44,301
Total trajectories (with all opus)
44,866
Categories
2,122
Avg turns per trajectory
1.2
Source dataset
GeneralThought… See the full description on the dataset page: https://huggingface.co/datasets/mai-ll/ipw-sft-trajectories.belebele-fi-filtered-sft
Dataset Card for Finnish-NLP/benebele
Creation process
Finnish subset loaded from facebook/belebele
distill_r1_110k_sft_modifiedBorrowed from https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT
Fix the <image> placeholder issue, which will cause error during training:
raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.")
