datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
novel-agent-sft-dataset
All Novel Can Be Galgame — 完整数据集
中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。
项目地址:https://github.com/lin1753/novel2galgame
训练代码仓库:https://github.com/lin1753/novel-agent
数据规模
目录
文件数
大小
说明
training/
52
689 MB
训练用 SFT 数据 (JSONL)
raw-books/
671
327 MB
669 本原始小说
processed/
39,842
1.2 GB
按章节预处理文本
annotations/
1,626
1 MB
原始标注文件
合计
42,191
2.2 GB
目录结构
datasets/
├── training/
│ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.zenyx-v2-SFT-dataset
Zenyx V2 — Raw SFT Dataset Collection
This is the unified raw dataset collection used for training Zenyx V2,
a custom large language model built from scratch with a novel architecture.
Dataset Sources
Dataset
Rows
Category
nemotron_sft_code
10,108,883
Code
nemotron_sft_math
22,066,397
Math
nemotron_sft_science
708,920
Science
nemotron_sft_chat
39,792
Chat
nemotron_sft_safety
31,426
Safety
nemotron_rl
56,339
Instruction Following (RL)… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v2-SFT-dataset.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.dromedario-3-sft-dataset
🐪 Dataset Card for Dromedario 3
📋 Dataset Summary
Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.Turkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.Agentic-Chain-of-Thought-Coding-SFT-Dataset
🤖 Agentic Coding CoT Dataset
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset.specialist-level_medical_knowledge_dataset_sft
specialist-level_medical_knowledge_dataset_sft
Dataset Summary
specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.kicad-netlist-sft-dataset
KiCad Netlist SFT Dataset
Training dataset for fine-tuning LLMs to generate valid KiCad electronic circuit netlists from natural language descriptions. Contains 100,179 examples with two complementary output formats:
Blog post: Teaching a Small LLM to Design Electronic Circuits: Fine-Tuning Qwen3-4B on 100K KiCad Netlists
Format
Examples
Description
SKiDL Python
100,179
Executable Python netlists in the messages assistant field
Structured JSON
100,179
Parallel… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/kicad-netlist-sft-dataset.Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
Llama-Nemotron-Post-Training-Dataset-SFT-math-FI
This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset.
The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model.
Translation Process
The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.full-structured-instruction-sft-dataset
Full Structured + Instruction SFT Corpus
Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data.
Dataset repo
mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11
Included files
train_full_sft.jsonl: full merged and shuffled SFT dataset
source_glaive.jsonl: processed Glaive subset
source_hermes.jsonl: processed Hermes subset
source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.SFT_Dataset
Pythagoras SFT Dataset
Project Page | GitHub | Paper
Data
Our training dataset consists of approximately 841K problems paired with Lean formal statements, formal proofs, and reasoning chains. We release a partial subset, which consists of 126K instances:
30K easy instances
49K medium instances
47K hard instances
Complete data will be released soon.
The complete explanation of the synthetic data generation pipeline can be found in Pythagoras-Prover: Advancing… See the full description on the dataset page: https://huggingface.co/datasets/Pythagoras-LM/SFT_Dataset.essential-level_medical_knowledge_dataset_sft
essential-level_medical_knowledge_dataset_sft
Dataset Summary
essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.cybersecurity-sft-dataset
Cybersecurity SFT Dataset
A curated dataset for training cybersecurity-focused code models with structured JSON output capability.
Dataset Composition
Source
Count
Percentage
Description
CVE Records
10,000
50.0%
Multi-turn CVE vulnerability analysis
OpenCodeReasoning (NVIDIA)
5,000
25.0%
Chain-of-thought code reasoning
Code-Feedback
5,000
25.0%
Multi-turn code debugging and refinement
Synthetic Security (JSON)
5
<0.1%
JSON-structured CVE, MITRE ATT&CK… See the full description on the dataset page: https://huggingface.co/datasets/moro72842/cybersecurity-sft-dataset.ja-safety-sft-dataset
ja-safety-sft-dataset
日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。
A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below.
概要
株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。
関連モデル
本サンプルの元データを用いて以下のモデルを安全性チューニングしました。
APTO-001/Qwen3.5-27B-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF)
APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.sft_alfworld_trajectory_dataset_v5
ALFWorld Trajectory Dataset
Overview
This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model.
Key Approach
Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples).
Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v5.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.devops-sft-dataset
DevOps SFT Instruction Dataset
This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline.
Dataset Description
Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/jalpan04/devops-sft-dataset.tau2-sft-v4-dataset
tau2-sft-v4-dataset
Expert trajectories for training tool-calling agents on tau2-bench tasks.
Overview
This dataset contains 219 multi-turn trajectories generated by Qwen3-235B-A22B-Thinking acting as a teacher model on the tau2-bench evaluation framework.
Dataset Statistics
Domain
Traces
Telecom
59
Airline
49
Retail
111
Total
219
Format
Each trace follows the tool-first format required by tau2-bench:
{
"task_id":… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/tau2-sft-v4-dataset.fire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.unified-sft-dataset
Loading
from datasets import load_dataset
ds = load_dataset("himalaya-ai/unified-sft-dataset")
Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1
🤖 Agentic Coding CoT Dataset v1.1
A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities.
📋 Dataset Description
This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns.
🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.PACE-SFT-DATASETS
PACE-SFT: Plot-Aware Continuation Evaluator — SFT Dataset
简介
PACE-SFT 是一个面向 AI 视频短剧 场景的中文剧情数据集,用于训练具备剧情续写、剧情分析和剧情质量评估能力的大语言模型。数据通过 DeepSeek-V3/R1 蒸馏生成,覆盖 10 种主流短剧类型。
本数据集是 PACE(Plot-Aware Continuation Evaluator)项目的 SFT 阶段训练数据,目标是为后续训练剧情续写质量评估的 Reward Model 打下基础。
应用场景
AI 视频剧情自动续写
多候选剧情质量排序
剧情连贯性和世界观一致性评估
短剧内容生产流水线中的质量把控
任务类型
任务
说明
占比
continuation
基于 IP 世界观设定和前文续写下一集剧情
~50%
analysis
对剧情进行结构化要素拆解(冲突/角色/悬念等)
~25%
evaluation… See the full description on the dataset page: https://huggingface.co/datasets/SuperYuanAI/PACE-SFT-DATASETS.typakos-sft-dataset
Typakos SFT Dataset
A bilingual (Greek/English) instruction-tuning dataset used to supervised-fine-tune alexliap/typakos-140m-base into alexliap/typakos-140m-it. It mixes 7 filtered subsets drawn from 4 upstream Hub sources into a single shuffled pool of 936,491 train and 104,055 validation conversations, each token-bounded to fit the model's 2048-token context length.
The full construction pipeline, code, and configs live in scripts/typakos_140m/ on GitHub.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/typakos-sft-dataset.SFT-Dataset
SFT-Dataset
A curated, medium-scale mixture designed to push a base model toward two things at once: stronger step-by-step reasoning (math, science, code) and reliable instruction following (format, language, and task constraints).
Quantities are chosen to stay trainable on modest GPU budgets while keeping signal density high—useful as a standalone SFT stage or as a clean warm start before reinforcement learning.
Evidence: benchmarks on a model trained on this mixture… See the full description on the dataset page: https://huggingface.co/datasets/SeaFill2025/SFT-Dataset.counter-sft-01-dataset
Counter-SFT-01
A synthetic conversational dataset for supervised fine-tuning on a
constrained counter-planning task.
The model must move a counter from start to target using increments
of 1, 2, or 3, with at most five increments.
Required response format:
<counter_plan>{"increments":[3,3,2],"final":12}</counter_plan>
Splits
Split
Rows
Start range
Templates
micro_train
32
0-10
A
train
480
0-59
A, B, C
validation
90
60-69
A, B, C
test
90
70-79
A, B… See the full description on the dataset page: https://huggingface.co/datasets/ishagarg1103/counter-sft-01-dataset.Bilingual-SFT-Dataset
Bilingual-SFT-Dataset
This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities.
Attributes:
Language(s): English, Pashto
License: apache-2.0
Size: 200,000 entries
Format: JSONL
Source: iPashto.ai
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.hindi-novel-sft-dataset
📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस)
यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है।
🌟 प्रमुख विशेषताएँ (Key Highlights)
10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ।
100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.AMALIA-LLM-0626-SFT-Dataset
AMALIA LLM Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training.
Base Data Mix
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
amalia-llm/persona_math
63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.
