CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikuhhn1239 /novel-agent-sft-dataset All Novel Can Be Galgame — 完整数据集 中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。 项目地址:https://github.com/lin1753/novel2galgame 训练代码仓库:https://github.com/lin1753/novel-agent 数据规模 目录 文件数 大小 说明 training/ 52 689 MB 训练用 SFT 数据 (JSONL) raw-books/ 671 327 MB 669 本原始小说 processed/ 39,842 1.2 GB 按章节预处理文本 annotations/ 1,626 1 MB 原始标注文件 合计 42,191 2.2 GB 目录结构 datasets/ ├── training/ │ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.texttext-generation10K<n<100K7 likes4.9k downloads3mo agoHugging Face02nvidia /Cosmos-Reason1-SFT-Dataset Dataset Description: The data format is a pair of video and text annotations. We summarize the data and annotations in Table 4 (SFT), Table 5 (RL), and Table 6 (Benchmark) of the Cosmos-Reason1 paper. ​​ We release the annotations for embodied reasoning tasks for BridgeDatav2, RoboVQA, Agibot, HoloAssist, AV, and the videos for the RoboVQA and AV datasets. We additionally release the annotations and videos for the RoboFail dataset for benchmarks. By releasing the dataset, NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Cosmos-Reason1-SFT-Dataset.textvisual-question-answering1M<n<10M31 likes1.9k downloads1y agoHugging Face03Logics-MLLM /Logics-STEM-SFT-Dataset-Open-1.6M Logics-STEM-SFT-Dataset-2.2M 📰 News [2026.01.05]🔥 Release of our Techinical Report. [2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M. Overview What is this dataset? Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.text1M<n<10M33 likes849 downloads8mo agoHugging Face04AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes426 downloads11mo agoHugging Face05AlicanKiraz0 /Agentic-Chain-of-Thought-Coding-SFT-Dataset 🤖 Agentic Coding CoT Dataset A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️ Assistant Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset.texttext-generationn<1K76 likes343 downloads10mo agoHugging Face06C3DS /cards_sft_dataset CARDS SFT — Climate Contrarian Discourse This is the dataset used to train the CARDS models released under C3DS (e.g. CARDS-Qwen3.6-27B, CARDS-Qwen3.5-{4B,9B,27B} and their FP8 / GGUF variants). It contains the supervised fine-tuning data and held-out evaluation splits for the hierarchical climate-discourse claim classifier from: Coan, T.G., Malla, R., Nanko, M.O., Kattrup, W., Roberts, J.T., Cook, J., Boussalis, C. Large language model reveals an increase in climate contrarian… See the full description on the dataset page: https://huggingface.co/datasets/C3DS/cards_sft_dataset.texttext-classification1K<n<10K1 likes294 downloads5mo agoHugging Face07snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes291 downloads1mo agoHugging Face08JunxiongWang /sftdatasetThis is the dataset used in paper, The Mamba in the Llama: Distilling and Accelerating Hybrid Models. @article{junxiongdaniele2024mambainllama, title = {The Mamba in the Llama: Distilling and Accelerating Hybrid Models}, author = {Junxiong Wang and Daniele Paliotta and Avner May and Alexander M. Rush and Tri Dao}, journal = {arXiv preprint arXiv:2408.15237}, year = {2024} } We collect and reformat dataset from those sources. https://huggingface.co/datasets/teknium/OpenHermes-2.5… See the full description on the dataset page: https://huggingface.co/datasets/JunxiongWang/sftdataset.text10M<n<100M2 likes226 downloads2y agoHugging Face09Logics-MLLM /Logics-STEM-SFT-Dataset-Open-5.3Mtext1M<n<10M4 likes222 downloads8mo agoHugging Face10JunxiongWang /sftdatasetv3This is the dataset used in paper, The Mamba in the Llama: Distilling and Accelerating Hybrid Models. @article{junxiongdaniele2024mambainllama, title = {The Mamba in the Llama: Distilling and Accelerating Hybrid Models}, author = {Junxiong Wang and Daniele Paliotta and Avner May and Alexander M. Rush and Tri Dao}, journal = {arXiv preprint arXiv:2408.15237}, year = {2024} } We collect and reformat dataset from those sources. https://huggingface.co/datasets/teknium/OpenHermes-2.5… See the full description on the dataset page: https://huggingface.co/datasets/JunxiongWang/sftdatasetv3.text10M<n<100M1 likes202 downloads2y agoHugging Face11PersonalAILab /AFM-WebAgent-SFT-Dataset Data Introduction This dataset serves as the core training data for Agent Foundation Models (AFMs), specifically designed to elicit end-to-end multi-agent reasoning capabilities in large language models. Built on the novel "Chain-of-Agents (CoA)" paradigm, the dataset leverages a multi-agent distillation framework to transform collaboration processes from state-of-the-art multi-agent systems into trajectory data suitable for supervised fine-tuning (SFT), simulating dynamic… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/AFM-WebAgent-SFT-Dataset.text1K<n<10K10 likes191 downloads1y agoHugging Face12mdonigian /full-structured-instruction-sft-dataset Full Structured + Instruction SFT Corpus Unified SFT training corpus built from Glaive, Hermes, UltraChat, and synthetic structured-output data. Dataset repo mdonigian/full-structured-instruction-sft-datasetRelease date: 2026-03-11 Included files train_full_sft.jsonl: full merged and shuffled SFT dataset source_glaive.jsonl: processed Glaive subset source_hermes.jsonl: processed Hermes subset source_ultrachat.jsonl: processed UltraChat subset… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/full-structured-instruction-sft-dataset.texttext-generation10K<n<100K0 likes189 downloads7mo agoHugging Face13Pythagoras-LM /SFT_Dataset Pythagoras SFT Dataset Project Page | GitHub | Paper Data Our training dataset consists of approximately 841K problems paired with Lean formal statements, formal proofs, and reasoning chains. We release a partial subset, which consists of 126K instances: 30K easy instances 49K medium instances 47K hard instances Complete data will be released soon. The complete explanation of the synthetic data generation pipeline can be found in Pythagoras-Prover: Advancing… See the full description on the dataset page: https://huggingface.co/datasets/Pythagoras-LM/SFT_Dataset.texttext-generation100K<n<1M9 likes137 downloads3mo agoHugging Face14AlicanKiraz0 /Turkish-Finance-SFT-Dataset 🇹🇷 Turkish Finance SFT Dataset Türkçe Finans Alanına Özel Supervised Fine-Tuning (SFT) Dataseti 📋 Dataset Özeti Bu dataset, Türkçe finans asistanı LLM'lerin eğitimi için özel olarak tasarlanmış, kapsamlı bir Supervised Fine-Tuning (SFT) veri setidir. Kripto para, borsa, teknik analiz, temel analiz, risk yönetimi ve finansal regülasyonlar dahil olmak üzere geniş bir yelpazede yaklaşık 10 milyon token boyutunda soru-cevap çifti verisi içermektedir. Dataset, hem… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-Finance-SFT-Dataset.textquestion-answering1K<n<10K63 likes135 downloads7mo agoHugging Face15snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes135 downloads4d agoHugging Face16zhendongnvidia /qwen3-tool-calling-sft-dataset Tool Calling Dataset for Fine-Tuning High-quality tool calling dataset with consistent schema for supervised fine-tuning. Dataset Description This dataset contains 11 high-quality single-turn tool calling conversations in standard OpenAI chat completion format. Features ✅ Schema Consistent: All parameter types normalized across records ✅ Quality Filtered: GPT-4o-mini evaluated (score ≥ 7.0/10) ✅ OpenAI Compatible: Ready for direct use with OpenAI fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/zhendongnvidia/qwen3-tool-calling-sft-dataset.textn<1K0 likes134 downloads1y agoHugging Face17realrick /ch05-sft-dataset 金融领域 SFT 微调数据集 数据集概览 总条数: 3,144 格式: Alpaca (instruction / input / output) 语言: 中文 覆盖领域: 12个金融子领域 指令类型: 9种 覆盖领域 基础金融概念 - 货币、利率、通胀、GDP、CPI等 股票投资 - PE、PB、K线、技术分析、基本面等 基金理财 - 货币基金、债券基金、指数基金、ETF等 保险知识 - 重疾险、医疗险、寿险、年金险、核保理赔等 银行业务 - 存款、贷款、信用卡、大额存单、数字人民币等 个人理财 - 资产配置、复利、养老规划、教育金等 宏观经济 - 货币政策、财政政策、经济周期、国际贸易等 财务分析 - 三大报表、杜邦分析、财务比率等 风险管理 - VaR、对冲、分散化、压力测试等 税收知识 - 个税、企业所得税、增值税、印花税等 区块链与数字货币 - 比特币、以太坊、DeFi、智能合约等 金融法规 - 证券法、基金法、反洗钱、资管新规等 指令类型分布… See the full description on the dataset page: https://huggingface.co/datasets/realrick/ch05-sft-dataset.text1K<n<10K0 likes114 downloads3mo agoHugging Face18APTO-001 /ja-safety-sft-dataset ja-safety-sft-dataset 日本語LLMの安全性チューニング用 SFT データセットのサンプル (500件) です。 A 500-item sample of the SFT dataset used to safety-tune APTO's Japanese LLMs. English version is provided below. 概要 株式会社APTOが大規模言語モデル(LLM)の安全性向上のために作成した約18,000件の日本語安全性学習データから、比率を維持して抽出したサンプルです。本サンプルでデータの構造と品質を確認できます。 関連モデル 本サンプルの元データを用いて以下のモデルを安全性チューニングしました。 APTO-001/Qwen3.5-27B-SafetyTuned (GGUF) APTO-001/Qwen3.5-9B-Base-SafetyTuned (GGUF) APTO-001/Qwen3.5-9B-SafetyTuned (GGUF)… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/ja-safety-sft-dataset.texttext-generationn<1K0 likes113 downloads4mo agoHugging Face19u-10bei /sft_alfworld_trajectory_dataset_v5 ALFWorld Trajectory Dataset Overview This is a synthetic SFT (Supervised Fine-Tuning) dataset designed for agent training in ALFWorld-compatible environments. The dataset programmatically generates expert trajectories without requiring an actual ALFWorld environment or a large language model. Key Approach Template-based Simulation: Lightweight simulator based on published ALFWorld information (papers, ReAct prompt examples). Subgoal Decomposition: Rule-based… See the full description on the dataset page: https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v5.texttext-generation1K<n<10K0 likes109 downloads8mo agoHugging Face20jalpan04 /devops-sft-dataset DevOps SFT Instruction Dataset This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline. Dataset Description Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/jalpan04/devops-sft-dataset.texttext-generation1K<n<10K0 likes104 downloads3mo agoHugging Face21Jarrodbarnes /tau2-sft-v4-dataset tau2-sft-v4-dataset Expert trajectories for training tool-calling agents on tau2-bench tasks. Overview This dataset contains 219 multi-turn trajectories generated by Qwen3-235B-A22B-Thinking acting as a teacher model on the tau2-bench evaluation framework. Dataset Statistics Domain Traces Telecom 59 Airline 49 Retail 111 Total 219 Format Each trace follows the tool-first format required by tau2-bench: { "task_id":… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/tau2-sft-v4-dataset.texttext-generationn<1K0 likes101 downloads10mo agoHugging Face22sdzjoy /fire-safety-sft-dataset Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集 Overview / 概述 A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.textquestion-answering10K<n<100K2 likes99 downloads6mo agoHugging Face23AlicanKiraz0 /Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1 🤖 Agentic Coding CoT Dataset v1.1 A high-quality supervised fine-tuning (SFT) dataset for training agentic coding assistants with Chain-of-Thought reasoning capabilities. 📋 Dataset Description This dataset was created by processing and distilling ~20GB of GitHub crawl data using Minimax-M2 & MiniMax M2.1 to generate structured, reasoning-rich coding examples. Each sample demonstrates systematic problem-solving with explicit tool usage patterns. 🏗️ Assistant… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Agentic-Chain-of-Thought-Coding-SFT-Dataset-v1.1.texttext-generation1K<n<10K14 likes93 downloads9mo agoHugging Face24SuperYuanAI /PACE-SFT-DATASETS PACE-SFT: Plot-Aware Continuation Evaluator — SFT Dataset 简介 PACE-SFT 是一个面向 AI 视频短剧 场景的中文剧情数据集,用于训练具备剧情续写、剧情分析和剧情质量评估能力的大语言模型。数据通过 DeepSeek-V3/R1 蒸馏生成,覆盖 10 种主流短剧类型。 本数据集是 PACE(Plot-Aware Continuation Evaluator)项目的 SFT 阶段训练数据,目标是为后续训练剧情续写质量评估的 Reward Model 打下基础。 应用场景 AI 视频剧情自动续写 多候选剧情质量排序 剧情连贯性和世界观一致性评估 短剧内容生产流水线中的质量把控 任务类型 任务 说明 占比 continuation 基于 IP 世界观设定和前文续写下一集剧情 ~50% analysis 对剧情进行结构化要素拆解(冲突/角色/悬念等) ~25% evaluation… See the full description on the dataset page: https://huggingface.co/datasets/SuperYuanAI/PACE-SFT-DATASETS.texttext-generation1K<n<10K1 likes88 downloads7mo agoHugging Face25himalaya-ai /nepali-sft-datasettext1M<n<10M2 likes80 downloads6mo agoHugging Face26nassimjp /Bilingual-SFT-Dataset Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.texttext-generation100K<n<1M0 likes74 downloads1mo agoHugging Face27Omarrran /Sample_kashmiri_sft_dataset Dataset Details: This dataset is a sample of a Kashmiri SFT Q&A-based dataset. Currently, scaling it requires GPU compute. If you have the necessary compute resources, we can collaborate to scale this dataset further. please connect me at "hnm.cs.ai@outlook.com" textn<1K0 likes73 downloads6mo agoHugging Face28zhendongnvidia /qwen3-tool-calling-sft-dataset-1k Tool Calling Dataset for Fine-Tuning High-quality tool calling dataset for supervised fine-tuning. Dataset Description This dataset contains 847 high-quality single-turn tool calling conversations in standard OpenAI chat completion format, optimized for supervised fine-tuning. Features ✅ High Quality: GPT-4o-mini evaluated (score ≥ 6.0/10) ✅ SFT Ready: Truncated to user+assistant pairs for supervised fine-tuning ✅ OpenAI Compatible: Ready for direct use with… See the full description on the dataset page: https://huggingface.co/datasets/zhendongnvidia/qwen3-tool-calling-sft-dataset-1k.textn<1K1 likes72 downloads1y agoHugging Face29vikasaivyas /hindi-novel-sft-dataset 📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस) यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है। 🌟 प्रमुख विशेषताएँ (Key Highlights) 10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ। 100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.texttext-generationn<1K0 likes72 downloads16d agoHugging Face30PersonalAILab /O-Researcher-SFT-Dataset Data Introduction This dataset serves as the core training data for O-Researcher, specifically designed to elicit end-to-end, multi-turn, multi-tool deep research capabilities in large language models. Built on the Multi-Agent Data Synthesis paradigm, the dataset leverages collaborative AI agents to simulate complex tool-integrated reasoning, transforming multi-agent research workflows into trajectory data suitable for supervised fine-tuning (SFT), enabling dynamic web search, page… See the full description on the dataset page: https://huggingface.co/datasets/PersonalAILab/O-Researcher-SFT-Dataset.text1K<n<10K2 likes71 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.