CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4.5k downloads27d agoHugging Face02Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M6 likes1.9k downloads2mo agoHugging Face03Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M0 likes546 downloads26d agoHugging Face04AlicanKiraz0 /Turkish-SFT-Dataset-v1.0 Turkish-SFT-Dataset-v1.01 Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci 🔎 Özet Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.texttext-classification1K<n<10K52 likes426 downloads11mo agoHugging Face05snuh /specialist-level_medical_knowledge_dataset_sft specialist-level_medical_knowledge_dataset_sft Dataset Summary specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K1 likes291 downloads1mo agoHugging Face06AlicanKiraz0 /Turkish-Finance-SFT-Dataset 🇹🇷 Turkish Finance SFT Dataset Türkçe Finans Alanına Özel Supervised Fine-Tuning (SFT) Dataseti 📋 Dataset Özeti Bu dataset, Türkçe finans asistanı LLM'lerin eğitimi için özel olarak tasarlanmış, kapsamlı bir Supervised Fine-Tuning (SFT) veri setidir. Kripto para, borsa, teknik analiz, temel analiz, risk yönetimi ve finansal regülasyonlar dahil olmak üzere geniş bir yelpazede yaklaşık 10 milyon token boyutunda soru-cevap çifti verisi içermektedir. Dataset, hem… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-Finance-SFT-Dataset.textquestion-answering1K<n<10K63 likes135 downloads7mo agoHugging Face07snuh /essential-level_medical_knowledge_dataset_sft essential-level_medical_knowledge_dataset_sft Dataset Summary essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH. This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub. It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.textquestion-answering10K<n<100K0 likes135 downloads4d agoHugging Face08TypeSafeAI /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M1 likes107 downloads2d agoHugging Face09jalpan04 /devops-sft-dataset DevOps SFT Instruction Dataset This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline. Dataset Description Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/jalpan04/devops-sft-dataset.texttext-generation1K<n<10K0 likes104 downloads3mo agoHugging Face10sdzjoy /fire-safety-sft-dataset Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集 Overview / 概述 A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.textquestion-answering10K<n<100K2 likes99 downloads6mo agoHugging Face11himalaya-ai /unified-sft-dataset Loading from datasets import load_dataset ds = load_dataset("himalaya-ai/unified-sft-dataset") texttext-generation100K<n<1M1 likes98 downloads28d agoHugging Face12somasekhar-dev /nexttoken-model-1-dataset-sft NextToken Model 1 SFT dataset (v4) Grounded multilingual QA dataset for fine-tuning somasekhar-dev/NextToken-model-1 on the Indian government-schemes / banking-financial domain. Generated by a pipeline (chunk source docs -> generate questions -> generate grounded answers -> validate/assemble) using a local LLM generator, from 846 scheme/product source documents across 57 schemes/products, chunked into 1,445 passages. v4 vs v3: v3 merged in a second batch (6,664 rows) without… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-1-dataset-sft.question-answering0 likes76 downloads6d agoHugging Face13nassimjp /Bilingual-SFT-Dataset Bilingual-SFT-Dataset This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities. Attributes: Language(s): English, Pashto License: apache-2.0 Size: 200,000 entries Format: JSONL Source: iPashto.ai Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.texttext-generation100K<n<1M0 likes74 downloads1mo agoHugging Face14amalia-llm /AMALIA-LLM-0626-SFT-Dataset AMALIA LLM Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training. Base Data Mix This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count amalia-llm/persona_math 63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.texttext-generation1M<n<10M2 likes68 downloads3mo agoHugging Face15cloudfrm-site /unified-sft-dataset Loading from datasets import load_dataset ds = load_dataset("himalaya-ai/unified-sft-dataset") texttext-generation100K<n<1M0 likes66 downloads24d agoHugging Face16SharathReddy /Indian-Legal-SFT-Dataset Vidhaan: High-Density Indian Legal Instruction Dataset Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets. 🛠 Dataset Structure & Format Primary File: vidhaan_training_v1.jsonl Format: JSON Lines (JSONL) Schema: - instruction: (String) A precise legal query. context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.textquestion-answering10K<n<100K0 likes61 downloads6mo agoHugging Face17justinthelaw /Resume-Cover-Letter-SFT-Dataset justinthelaw/Resume-Cover-Letter-SFT-Dataset A Supervised Fine-Tuning (SFT) dataset generated from Justin's resume for fine-tuning language models to answer questions about professional background, skills, and experience. This dataset consists of synthetically generated QA pairs. Dataset Statistics Total Samples: ~9600 (estimated, with 3x variations per unique question) Train Split: 90% Validation Split: 9% Samples per Category: 1200 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/justinthelaw/Resume-Cover-Letter-SFT-Dataset.text-generation1K<n<10K0 likes59 downloads7mo agoHugging Face18AayushP418 /finlora-sft-dataset FinLoRA SFT Dataset ChatML-formatted instruction-tuning dataset for SEC 10-K financial document question answering. Used to train finlora-sft-v2-phi35 via QLoRA supervised fine-tuning on Phi-3.5-mini-instruct. Dataset summary Split Examples Size train 31,320 ~99 MB validation 3,479 ~11 MB Total: 34,799 examples Sources EDGAR (28,548 examples) — 125 SEC 10-K annual reports across 25 S&P 500 companies, 5 sectors (Technology, Finance… See the full description on the dataset page: https://huggingface.co/datasets/AayushP418/finlora-sft-dataset.textquestion-answering10K<n<100K0 likes49 downloads3mo agoHugging Face19raghu298 /ml-interview-sft-dataset ML/AI Interview Coach — SFT Dataset A curated dataset of 566 high-quality Q&A pairs covering ML, Deep Learning, NLP, LLMs, RAG, Vector Databases, LangChain, Agentic AI, MLOps, and more — designed for fine-tuning an ML Interview Coach model. Dataset Summary Stat Value Total Q&A pairs 566 Unique topics 75 Format ChatML (system + user + assistant) Language English Avg answer length ~800 tokens Sources 15+ interview prep documents + hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/raghu298/ml-interview-sft-dataset.textquestion-answeringn<1K0 likes45 downloads5mo agoHugging Face20Phonsiri /legal-chat-sft-dataset Thai Legal Chat SFT Dataset (CoT & Hybrid RAG) ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation) Dataset Summary ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.texttext-generation10K<n<100K1 likes43 downloads5mo agoHugging Face21Somtharu181coder /Legal_domain_ocr_extracted_Nepali_sft_dataset Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083 Dataset Summary This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument: सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३ (Software Development and Operation Committee (Formation) Order, 2083) The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.texttext-generationn<1K0 likes43 downloads16d agoHugging Face22ChaoticEconomist /Classical-Mechanics-Equations-Dataset_SFT-or-LoRA Classical Mechanics Equations Dataset (SFT / LoRA Ready) A structured dataset of 64 classical mechanics equations from Newtonian, Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning rows across three task types: equation explanation, Q&A, and derivation. Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and equation understanding tasks. Overview Property Value Domain Classical Mechanics (Physics) Total rows 448 Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.texttext-generationn<1K0 likes41 downloads5mo agoHugging Face23SerFabio89 /italian-open-sft-chat-dataset Italian Open SFT Chat Dataset An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records. This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.texttext-generation10K<n<100K0 likes40 downloads4mo agoHugging Face24hozifa1 /faqih_sft_dataset 💎 MAFQA: Perfected Multi-Hop Arabic Fatwa QA Dataset (388 Samples) مجموعة بيانات الاستدلال الفقهي المركب ومتعدد الخطوات (جامعة الملك سعود / MDPI 2026) 100% Curated & Unabridged MAFQA Multi-Hop Dataset تم تنقيح وتدقيق البيانات بالكامل: 1. إزالة جميع التقطيعات النصية وإيراد النصوص والأدلة كاملة دون بتر. 2. تفعيل خطوة التركيب والترجيح النهائي (الخطوة 4) بربط استدلالي حقيقي بين المسائل الفرعية. 3. تصحيح الأخطاء المطبعية في دلالات الحل والحرمة.… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/faqih_sft_dataset.tabulartext-generation1K<n<10K0 likes38 downloads1mo agoHugging Face25fenyo /L40S-MonEspaceSante-SFT-dataset Mon Espace Santé — Données SFT (Q/R) Paires question/réponse en français pour l'étape SFT (instruction-following / format) du modèle fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT. Conformément à Gekhman et al. (arXiv:2405.05904), la connaissance est injectée au CPT (cf. corpus CPT) ; le SFT ne sert qu'à restaurer le format Q/R, pas à mémoriser. Composition (2 771 paires, après décontamination) Toutes les paires sont 1-hop des 88 faits réels : real — les 88 faits… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-SFT-dataset.textquestion-answering1K<n<10K0 likes35 downloads4mo agoHugging Face26likhitjuttada /finance-reasoning-sft-dataset Personal Finance Reasoning Dataset A synthetic instruction-tuning dataset designed to teach language models to reason through personal finance and investing decisions using the mental frameworks from classic books in the genre. The goal is not recall of book content but principled reasoning: the model should apply frameworks to novel situations it has never seen. Source Books Principles were extracted from the following books: The Psychology of Money — Morgan Housel Rich… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/finance-reasoning-sft-dataset.texttext-generationn<1K0 likes32 downloads5mo agoHugging Face27Minuri /sinhala-sft-dataset Sinhala Supervised Fine-Tuning Dataset A merged Sinhala instruction-following dataset of 213,703 pairs, used for Supervised Fine-Tuning (SFT) of continually pretrained LLaMA 3.2 1B variants. Constructed as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This dataset merges three existing Sinhala instruction datasets into a unified resource for SFT. It follows the standard Alpaca-style instruction–input–output format and covers a… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-sft-dataset.texttext-generation100K<n<1M0 likes31 downloads6mo agoHugging Face28amalia-llm /AMALIA-LLM-1225-SFT-Dataset AMALIA LLM 1225 Supervised Finetuning Dataset Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model version released in December 2025. This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table: Dataset Count Persona-PT Instruction Following 9,084 Persona-EN Instruction Following 34,704 Persona Nemotron Instruction Following 4,483… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-1225-SFT-Dataset.texttext-generation1M<n<10M0 likes31 downloads3mo agoHugging Face29kshitijthakkar /loggenix-stage4-sft-dataset LogGenix MoE Stage 4 SFT Dataset This dataset is designed for Stage 4 SFT training of the LogGenix MoE model, focusing on: Coherence Recovery - Fix Stage 2 damage, restore general capabilities Tool Calling - TraceVerse MCP tool invocation Trace Analysis - OpenTelemetry span analysis (synthetic) GPU Metrics Analysis - GPU monitoring and analysis Prompt Optimization - Help users write better prompts Dataset Summary Total Samples: 226,725 Total Tokens: 574,354,578… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/loggenix-stage4-sft-dataset.texttext-generation100K<n<1M0 likes30 downloads8mo agoHugging Face30ZhangYuchi /modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO OminiGAIA-DPO-data This dataset contains the final DPO training pairs used to train ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO. The pairs are derived from 7B-native rollouts on OmniGAIA train questions: Roll out the SFT model on answer-hidden train inputs. Audit each rollout with Gemini using the private reference answer and annotated solution. Locate the first erroneous assistant sub-step. Convert the corrected prefix (tau_win) and the original erroneous prefix… See the full description on the dataset page: https://huggingface.co/datasets/ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO.reinforcement-learning1 likes29 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.