CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikuhhn1239 /novel-agent-sft-dataset All Novel Can Be Galgame — 完整数据集 中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。 项目地址:https://github.com/lin1753/novel2galgame 训练代码仓库:https://github.com/lin1753/novel-agent 数据规模 目录 文件数 大小 说明 training/ 52 689 MB 训练用 SFT 数据 (JSONL) raw-books/ 671 327 MB 669 本原始小说 processed/ 39,842 1.2 GB 按章节预处理文本 annotations/ 1,626 1 MB 原始标注文件 合计 42,191 2.2 GB 目录结构 datasets/ ├── training/ │ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.texttext-generation10K<n<100K7 likes4.9k downloads3mo agoHugging Face02SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4.5k downloads27d agoHugging Face03Arko007 /zenyx-v2-SFT-dataset Zenyx V2 — Raw SFT Dataset Collection This is the unified raw dataset collection used for training Zenyx V2, a custom large language model built from scratch with a novel architecture. Dataset Sources Dataset Rows Category nemotron_sft_code 10,108,883 Code nemotron_sft_math 22,066,397 Math nemotron_sft_science 708,920 Science nemotron_sft_chat 39,792 Chat nemotron_sft_safety 31,426 Safety nemotron_rl 56,339 Instruction Following (RL)… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/zenyx-v2-SFT-dataset.texttext-generation10M<n<100M2 likes2.4k downloads6mo agoHugging Face04Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M6 likes1.9k downloads2mo agoHugging Face05Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes546 downloads26d agoHugging Face06sapienzanlp /dromedario-3-sft-dataset 🐪 Dataset Card for Dromedario 3 📋 Dataset Summary Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.texttext-generation100K<n<1M9 likes502 downloads8d agoHugging Face07AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes426 downloads11mo agoHugging Face08AlicanKiraz0 /Agentic-Chain-of-Thought-Coding-SFT-Dataset 🤖 Agentic Coding CoT Dataset A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️ Assistant Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset.texttext-generationn<1K76 likes343 downloads10mo agoHugging Face09snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes291 downloads1mo agoHugging Face10AbijahKaj /kicad-netlist-sft-dataset KiCad Netlist SFT Dataset Training dataset for fine-tuning LLMs to generate valid KiCad electronic circuit netlists from natural language descriptions. Contains 100,179 examples with two complementary output formats: Blog post: Teaching a Small LLM to Design Electronic Circuits: Fine-Tuning Qwen3-4B on 100K KiCad Netlists Format Examples Description SKiDL Python 100,179 Executable Python netlists in the messages assistant field Structured JSON 100,179 Parallel… See the full description on the dataset page: https://huggingface.co/datasets/AbijahKaj/kicad-netlist-sft-dataset.texttext-generation100K<n<1M1 likes232 downloads6d agoHugging Face11LumiOpen /Llama-Nemotron-Post-Training-Dataset-SFT-math-FI Llama-Nemotron-Post-Training-Dataset-SFT-math-FI This dataset is a Finnish machine-translated version of the SFT/math split from the original nvidia/Llama-Nemotron-Post-Training-Dataset. The data was created by translating the original English math SFT subset into Finnish using the DeepSeek-V3 model. Translation Process The user prompt and the thinking traces were translated separately in two LLM requests. For traces, the <think> and </think> tokens were preserved… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/Llama-Nemotron-Post-Training-Dataset-SFT-math-FI.texttext-generation1M<n<10M1 likes229 downloads2mo agoHugging Face12mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes189 downloads7mo agoHugging Face13Pythagoras-LM /SFT_Dataset Pythagoras SFT Dataset Project Page | GitHub | Paper Data Our training dataset consists of approximately 841K problems paired with Lean formal statements, formal proofs, and reasoning chains. We release a partial subset, which consists of 126K instances: 30K easy instances 49K medium instances 47K hard instances Complete data will be released soon. The complete explanation of the synthetic data generation pipeline can be found in Pythagoras-Prover: Advancing… See the full description on the dataset page: https://huggingface.co/datasets/Pythagoras-LM/SFT_Dataset.texttext-generation100K<n<1M9 likes137 downloads3mo agoHugging Face14snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes135 downloads4d agoHugging Face15moro72842 /cybersecurity-sft-dataset Cybersecurity SFT Dataset A curated dataset for training cybersecurity-focused code models with structured JSON output capability. Dataset Composition Source Count Percentage Description CVE Records 10,000 50.0% Multi-turn CVE vulnerability analysis OpenCodeReasoning (NVIDIA) 5,000 25.0% Chain-of-thought code reasoning Code-Feedback 5,000 25.0% Multi-turn code debugging and refinement Synthetic Security (JSON) 5 <0.1% JSON-structured CVE, MITRE ATT&CK… See the full description on the dataset page: https://huggingface.co/datasets/moro72842/cybersecurity-sft-dataset.texttext-generation10K<n<100K1 likes123 downloads5mo agoHugging Face16APTO-001 /ja-safety-sft-dataset ja-safety-sft-dataset 日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。 A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below. 概要 株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。 関連モデル 本サンプルの元データを用いて以下のモデルを安全性チューニングしました。 APTO-001/Qwen3.5-27B-SafetyTuned (GGUF) APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF) APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.texttext-generationn<1K0 likes113 downloads4mo agoHugging Face17u-10bei /sft_alfworld_trajectory_dataset_v5 ALFWorld Trajectory Dataset Overview This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model. Key Approach Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples). Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v5.texttext-generation1K<n<10K0 likes109 downloads8mo agoHugging Face18TypeSafeAI /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M1 likes107 downloads2d agoHugging Face19jalpan04 /devops-sft-dataset DevOps SFT Instruction Dataset This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline. Dataset Description Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/jalpan04/devops-sft-dataset.texttext-generation1K<n<10K0 likes104 downloads3mo agoHugging Face20Jarrodbarnes /tau2-sft-v4-dataset tau2-sft-v4-dataset Expert trajectories for training tool-calling agents on tau2-bench tasks. Overview This dataset contains 219 multi-turn trajectories generated by Qwen3-235B-A22B-Thinking acting as a teacher model on the tau2-bench evaluation framework. Dataset Statistics Domain Traces Telecom 59 Airline 49 Retail 111 Total 219 Format Each trace follows the tool-first format required by tau2-bench: { "task_id":… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/tau2-sft-v4-dataset.texttext-generationn<1K0 likes101 downloads10mo agoHugging Face21sdzjoy /fire-safety-sft-dataset Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集 Overview / 概述 A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.textquestion-answering10K<n<100K2 likes99 downloads6mo agoHugging Face22himalaya-ai /unified-sft-dataset Loading from datasets import load_dataset ds = load_dataset("himalaya-ai/unified-sft-dataset") texttext-generation100K<n<1M1 likes98 downloads28d agoHugging Face23AlicanKiraz0 /Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1 🤖 Agentic Coding CoT Dataset v1.1 A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.texttext-generation1K<n<10K14 likes93 downloads9mo agoHugging Face24SuperYuanAI /PACE-SFT-DATASETS PACE-SFT: Plot-Aware Continuation Evaluator — SFT Dataset 简介 PACE-SFT 是一个面向 AI 视频短剧 场景的中文剧情数据集,用于训练具备剧情续写、剧情分析和剧情质量评估能力的大语言模型。数据通过 DeepSeek-V3/R1 蒸馏生成,覆盖 10 种主流短剧类型。 本数据集是 PACE(Plot-Aware Continuation Evaluator)项目的 SFT 阶段训练数据,目标是为后续训练剧情续写质量评估的 Reward Model 打下基础。 应用场景 AI 视频剧情自动续写 多候选剧情质量排序 剧情连贯性和世界观一致性评估 短剧内容生产流水线中的质量把控 任务类型 任务 说明 占比 continuation 基于 IP 世界观设定和前文续写下一集剧情 ~50% analysis 对剧情进行结构化要素拆解(冲突/角色/悬念等) ~25% evaluation… See the full description on the dataset page: https://huggingface.co/datasets/SuperYuanAI/PACE-SFT-DATASETS.texttext-generation1K<n<10K1 likes88 downloads7mo agoHugging Face25alexliap /typakos-sft-dataset Typakos SFT Dataset A bilingual (Greek/English) instruction-tuning dataset used to supervised-fine-tune alexliap/typakos-140m-base into alexliap/typakos-140m-it. It mixes 7 filtered subsets drawn from 4 upstream Hub sources into a single shuffled pool of 936,491 train and 104,055 validation conversations, each token-bounded to fit the model's 2048-token context length. The full construction pipeline, code, and configs live in scripts/typakos_140m/ on GitHub. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alexliap/typakos-sft-dataset.tabulartext-generation1M<n<10M0 likes85 downloads1mo agoHugging Face26SeaFill2025 /SFT-Dataset SFT-Dataset A curated, medium-scale mixture designed to push a base model toward two things at once: stronger step-by-step reasoning (math, science, code) and reliable instruction following (format, language, and task constraints). Quantities are chosen to stay trainable on modest GPU budgets while keeping signal density high—useful as a standalone SFT stage or as a clean warm start before reinforcement learning. Evidence: benchmarks on a model trained on this mixture… See the full description on the dataset page: https://huggingface.co/datasets/SeaFill2025/SFT-Dataset.texttext-generation10K<n<100K4 likes79 downloads6mo agoHugging Face27ishagarg1103 /counter-sft-01-dataset Counter-SFT-01 A synthetic conversational dataset for supervised fine-tuning on a constrained counter-planning task. The model must move a counter from start to target using increments of 1, 2, or 3, with at most five increments. Required response format: <counter_plan>{"increments":[3,3,2],"final":12}</counter_plan> Splits Split Rows Start range Templates micro_train 32 0-10 A train 480 0-59 A, B, C validation 90 60-69 A, B, C test 90 70-79 A, B… See the full description on the dataset page: https://huggingface.co/datasets/ishagarg1103/counter-sft-01-dataset.tabulartext-generationn<1K0 likes77 downloads18d agoHugging Face28nassimjp /Bilingual-SFT-Dataset Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.texttext-generation100K<n<1M0 likes74 downloads1mo agoHugging Face29vikasaivyas /hindi-novel-sft-dataset 📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस) यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है। 🌟 प्रमुख विशेषताएँ (Key Highlights) 10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ। 100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.texttext-generationn<1K0 likes72 downloads16d agoHugging Face30amalia-llm /AMALIA-LLM-0626-SFT-Dataset AMALIA LLM Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training. Base Data Mix This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count amalia-llm/persona_math 63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.texttext-generation1M<n<10M2 likes68 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.