datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.Turkish-SFT-Dataset-v1.0
Turkish-SFT-Dataset-v1.01
Repo: AlicanKiraz0/Turkish-SFT-Dataset-v1.0Sürüm: v1.01Lisans: MITBiçim: jsonl (kolonlar: system, user, assistant)Boyut: ~5500 satır ve satır başına 3.000–4.500 token/satır (≈ 20M+ token)Dil: Türkçe (tr)Görevler: talimat izleme, SFT, muhakeme, güvenli ret, uzun-bağlam ve araç kullanım bilinci
🔎 Özet
Bu veri kümesi, Türkçe Denetimli İnce Ayar (SFT) için tasarlanmış, yüksek kaliteli ve uzun çıktılar içeren örneklerden oluşur. İçerik 12 ana… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-SFT-Dataset-v1.0.specialist-level_medical_knowledge_dataset_sft
specialist-level_medical_knowledge_dataset_sft
Dataset Summary
specialist-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 13 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Specialized Medical Knowledge Data (전문 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/specialist-level_medical_knowledge_dataset_sft.Turkish-Finance-SFT-Dataset
🇹🇷 Turkish Finance SFT Dataset
Türkçe Finans Alanına Özel Supervised Fine-Tuning (SFT) Dataseti
📋 Dataset Özeti
Bu dataset, Türkçe finans asistanı LLM'lerin eğitimi için özel olarak tasarlanmış, kapsamlı bir Supervised Fine-Tuning (SFT) veri setidir. Kripto para, borsa, teknik analiz, temel analiz, risk yönetimi ve finansal regülasyonlar dahil olmak üzere geniş bir yelpazede yaklaşık 10 milyon token boyutunda soru-cevap çifti verisi içermektedir.
Dataset, hem… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Turkish-Finance-SFT-Dataset.essential-level_medical_knowledge_dataset_sft
essential-level_medical_knowledge_dataset_sft
Dataset Summary
essential-level_medical_knowledge_dataset_sft is an integrated collection of augmented SFT data across 4 distinct medical domains, developed by the Healthcare AI Research Institute (HARI) at SNUH.
This dataset is derived and augmented from the Essential Medical Knowledge Data (필수의료 의학지식 데이터) provided by AI-Hub.
It focuses exclusively on complex clinical scenarios generated using the "Add Constraints"… See the full description on the dataset page: https://huggingface.co/datasets/snuh/essential-level_medical_knowledge_dataset_sft.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/TypeSafeAI/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.devops-sft-dataset
DevOps SFT Instruction Dataset
This dataset contains 8,076 high-quality instruction-response pairs specifically generated for fine-tuning a DevOps domain-specialized language model. It was used in the Supervised Fine-Tuning (SFT) phase of the Ulysses model training pipeline.
Dataset Description
Instructions were generated using the Gemini API (gemini-2.0-flash) and Ollama (qwen2.5-coder:7b) by feeding chunks of official DevOps documentation and GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/jalpan04/devops-sft-dataset.fire-safety-sft-dataset
Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集
Overview / 概述
A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.unified-sft-dataset
Loading
from datasets import load_dataset
ds = load_dataset("himalaya-ai/unified-sft-dataset")
nexttoken-model-1-dataset-sft
NextToken Model 1 SFT dataset (v4)
Grounded multilingual QA dataset for fine-tuning
somasekhar-dev/NextToken-model-1
on the Indian government-schemes / banking-financial domain.
Generated by a pipeline (chunk source docs -> generate questions -> generate
grounded answers -> validate/assemble) using a local LLM generator, from
846 scheme/product source documents across 57 schemes/products, chunked
into 1,445 passages.
v4 vs v3: v3 merged in a second batch (6,664 rows) without… See the full description on the dataset page: https://huggingface.co/datasets/somasekhar-dev/nexttoken-model-1-dataset-sft.Bilingual-SFT-Dataset
Bilingual-SFT-Dataset
This dataset is a general-purpose bilingual Supervised Fine-Tuning (SFT) dataset designed for training Large Language Models (LLMs) to handle both English and Pashto languages effectively. It is structured to create robust multilingual models by maintaining English proficiency while building Pashto capabilities.
Attributes:
Language(s): English, Pashto
License: apache-2.0
Size: 200,000 entries
Format: JSONL
Source: iPashto.ai
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Bilingual-SFT-Dataset.AMALIA-LLM-0626-SFT-Dataset
AMALIA LLM Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model. This data mix includes both the mixes used in the base and ramp down phases of the SFT training.
Base Data Mix
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
amalia-llm/persona_math
63,731… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-0626-SFT-Dataset.unified-sft-dataset
Loading
from datasets import load_dataset
ds = load_dataset("himalaya-ai/unified-sft-dataset")
Indian-Legal-SFT-Dataset
Vidhaan: High-Density Indian Legal Instruction Dataset
Vidhaan is a comprehensive, high-precision instruction-tuning dataset containing 20,690 QA pairs derived from 113 Central Acts of India. It was built specifically to solve the "context-splitting" problem found in standard legal RAG datasets.
🛠 Dataset Structure & Format
Primary File: vidhaan_training_v1.jsonl
Format: JSON Lines (JSONL)
Schema: - instruction: (String) A precise legal query.
context: (String) The… See the full description on the dataset page: https://huggingface.co/datasets/SharathReddy/Indian-Legal-SFT-Dataset.Resume-Cover-Letter-SFT-Dataset
justinthelaw/Resume-Cover-Letter-SFT-Dataset
A Supervised Fine-Tuning (SFT) dataset generated from Justin's resume for fine-tuning language models to answer questions about professional background, skills, and experience. This dataset consists of synthetically generated QA pairs.
Dataset Statistics
Total Samples: ~9600 (estimated, with 3x variations per unique question)
Train Split: 90%
Validation Split: 9%
Samples per Category: 1200
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/justinthelaw/Resume-Cover-Letter-SFT-Dataset.finlora-sft-dataset
FinLoRA SFT Dataset
ChatML-formatted instruction-tuning dataset for SEC 10-K financial document question answering.
Used to train finlora-sft-v2-phi35
via QLoRA supervised fine-tuning on Phi-3.5-mini-instruct.
Dataset summary
Split
Examples
Size
train
31,320
~99 MB
validation
3,479
~11 MB
Total: 34,799 examples
Sources
EDGAR (28,548 examples) — 125 SEC 10-K annual reports across 25 S&P 500 companies,
5 sectors (Technology, Finance… See the full description on the dataset page: https://huggingface.co/datasets/AayushP418/finlora-sft-dataset.ml-interview-sft-dataset
ML/AI Interview Coach — SFT Dataset
A curated dataset of 566 high-quality Q&A pairs covering ML, Deep Learning, NLP, LLMs, RAG, Vector Databases, LangChain, Agentic AI, MLOps, and more — designed for fine-tuning an ML Interview Coach model.
Dataset Summary
Stat
Value
Total Q&A pairs
566
Unique topics
75
Format
ChatML (system + user + assistant)
Language
English
Avg answer length
~800 tokens
Sources
15+ interview prep documents + hand-crafted… See the full description on the dataset page: https://huggingface.co/datasets/raghu298/ml-interview-sft-dataset.legal-chat-sft-dataset
Thai Legal Chat SFT Dataset (CoT & Hybrid RAG)
ชุดข้อมูลสำหรับการทำ Instruction Fine-Tuning (SFT) เพื่อสร้าง AI ผู้ช่วยนักกฎหมายไทยที่มีความสามารถในการคิดวิเคราะห์แบบเป็นขั้นตอน (Chain-of-Thought) และมีความรู้กฎหมายที่ทันสมัยจากการใช้ Hybrid RAG (Retrieval-Augmented Generation)
Dataset Summary
ชุดข้อมูลนี้ถูกสร้างขึ้นแบบสังเคราะห์ (Synthetic Data) โดยใช้โมเดลภาษาขนาดใหญ่ (LLM) ตระกูล Qwen (27B+) บนขุมพลัง AMD MI300X ผ่านระบบ vLLM Monster Engine… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/legal-chat-sft-dataset.Legal_domain_ocr_extracted_Nepali_sft_dataset
Nepali Legal SFT Dataset — Software Development & Operation Committee Order, 2083
Dataset Summary
This dataset contains 31 single-turn instruction/response pairs in Nepali (Devanagari script), derived from a single Government of Nepal legal instrument:
सफ्टवेयर विकास तथा सञ्चालन समिति (गठन) आदेश, २०८३
(Software Development and Operation Committee (Formation) Order, 2083)
The order was issued by the Government of Nepal under Section 3 of the विकास समिति ऐन, २०१३… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/Legal_domain_ocr_extracted_Nepali_sft_dataset.Classical-Mechanics-Equations-Dataset_SFT-or-LoRA
Classical Mechanics Equations Dataset (SFT / LoRA Ready)
A structured dataset of 64 classical mechanics equations from Newtonian,
Lagrangian, and Hamiltonian mechanics, expanded into 448 instruction-tuning
rows across three task types: equation explanation, Q&A, and derivation.
Designed for fine-tuning LLMs on physics reasoning, STEM Q&A, and
equation understanding tasks.
Overview
Property
Value
Domain
Classical Mechanics (Physics)
Total rows
448
Train… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticEconomist/Classical-Mechanics-Equations-Dataset_SFT-or-LoRA.italian-open-sft-chat-dataset
Italian Open SFT Chat Dataset
An Italian-first, model-neutral synthetic SFT and chat dataset for fine-tuning Italian-capable LLMs. It targets instruction tuning, Italian chat behavior, structured output generation, JSON/YAML/CSV format following, coding assistance, safety refusals, multi-turn dialogue and reasoning-style final answers. This v0.1.0 package does not include long-context QA records.
This dataset is intended for users searching for an Italian instruction tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-open-sft-chat-dataset.faqih_sft_dataset
💎 MAFQA: Perfected Multi-Hop Arabic Fatwa QA Dataset (388 Samples)
مجموعة بيانات الاستدلال الفقهي المركب ومتعدد الخطوات (جامعة الملك سعود / MDPI 2026)
100% Curated & Unabridged MAFQA Multi-Hop Dataset
تم تنقيح وتدقيق البيانات بالكامل:
1. إزالة جميع التقطيعات النصية وإيراد النصوص والأدلة كاملة دون بتر.
2. تفعيل خطوة التركيب والترجيح النهائي (الخطوة 4) بربط استدلالي حقيقي بين المسائل الفرعية.
3. تصحيح الأخطاء المطبعية في دلالات الحل والحرمة.… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/faqih_sft_dataset.L40S-MonEspaceSante-SFT-dataset
Mon Espace Santé — Données SFT (Q/R)
Paires question/réponse en français pour l'étape SFT (instruction-following / format) du modèle
fenyo/L40S-Qwen3-8B-MonEspaceSante-CPT-SFT.
Conformément à Gekhman et al. (arXiv:2405.05904), la connaissance est injectée au CPT (cf.
corpus CPT) ; le SFT ne sert qu'à
restaurer le format Q/R, pas à mémoriser.
Composition (2 771 paires, après décontamination)
Toutes les paires sont 1-hop des 88 faits réels :
real — les 88 faits… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/L40S-MonEspaceSante-SFT-dataset.finance-reasoning-sft-dataset
Personal Finance Reasoning Dataset
A synthetic instruction-tuning dataset designed to teach language models to reason through personal finance and investing decisions using the mental frameworks from classic books in the genre. The goal is not recall of book content but principled reasoning: the model should apply frameworks to novel situations it has never seen.
Source Books
Principles were extracted from the following books:
The Psychology of Money — Morgan Housel
Rich… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/finance-reasoning-sft-dataset.sinhala-sft-dataset
Sinhala Supervised Fine-Tuning Dataset
A merged Sinhala instruction-following dataset of 213,703 pairs, used for Supervised Fine-Tuning (SFT) of continually pretrained LLaMA 3.2 1B variants. Constructed as part of a diversity-driven Sinhala language model adaptation study.
Dataset Description
This dataset merges three existing Sinhala instruction datasets into a unified resource for SFT. It follows the standard Alpaca-style instruction–input–output format and covers a… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-sft-dataset.AMALIA-LLM-1225-SFT-Dataset
AMALIA LLM 1225 Supervised Finetuning Dataset
Data mix used in the Supervised Finetuning stage of the post-training of the AMALIA model version released in December 2025.
This data mix contains off-the-shelf datasets and developed by the AMALIA team. The dataset counts are described in the following table:
Dataset
Count
Persona-PT Instruction Following
9,084
Persona-EN Instruction Following
34,704
Persona Nemotron Instruction Following
4,483… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-LLM-1225-SFT-Dataset.loggenix-stage4-sft-dataset
LogGenix MoE Stage 4 SFT Dataset
This dataset is designed for Stage 4 SFT training of the LogGenix MoE model, focusing on:
Coherence Recovery - Fix Stage 2 damage, restore general capabilities
Tool Calling - TraceVerse MCP tool invocation
Trace Analysis - OpenTelemetry span analysis (synthetic)
GPU Metrics Analysis - GPU monitoring and analysis
Prompt Optimization - Help users write better prompts
Dataset Summary
Total Samples: 226,725
Total Tokens: 574,354,578… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/loggenix-stage4-sft-dataset.modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO
OminiGAIA-DPO-data
This dataset contains the final DPO training pairs used to train
ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO.
The pairs are derived from 7B-native rollouts on OmniGAIA train questions:
Roll out the SFT model on answer-hidden train inputs.
Audit each rollout with Gemini using the private reference answer and annotated solution.
Locate the first erroneous assistant sub-step.
Convert the corrected prefix (tau_win) and the original erroneous prefix… See the full description on the dataset page: https://huggingface.co/datasets/ZhangYuchi/modelbest-Qwen-2.5-Omni-7B-SFT-with-DPO-dataset_for_DPO.
