datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-reasoning-sft-100k
Math Reasoning SFT (100K)
100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models.
Dataset Description
100,000 problems across 8 mathematical categories and 3 difficulty levels:
Categories
Category
Examples
Topics
word_problems
~23,100
Rate/time/distance, work problems, mixture, meeting/catch-up
arithmetic
~15,400
Percentages, profit/loss, ratios
geometry
~15,400
Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.tool-calling-english-100k
Tool Calling English (100K)
100,000 tool-calling conversations in OpenAI function calling format — the largest general English tool-use dataset for fine-tuning.
Motivation
Models trained without tool-calling examples struggle in agentic deployments. This dataset trains the full cycle: deciding when to call a tool, calling it with correct arguments, interpreting the result, and producing a grounded final response.
Dataset Description
100,000… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/tool-calling-english-100k.medical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.writing-quality-dpo-100k
Writing Quality DPO (100K)
100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect.
Motivation
Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.email-writing-sft-100k
Email Writing SFT (100K)
100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails.
Motivation
Email is the primary communication channel for most professional work, yet LLMs often produce emails that are:
Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.science-qa-sft-100k
Science QA SFT (100K)
100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty.
Motivation
Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.sn38-quality-gold-100k
SN38 quality prompt + gold continuation
Synthetic incomplete-sentence prompts with gold continuations for Bittensor
subnet 38. 13 categories, 13 items per category per call,
temperature 1.0.
Each row:
category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning, math_reasoning, truthfulness, pronoun_resolution, paraphrase_detection, word_sense_disambiguation
prompt:… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-100k.technical-writing-sft-100k
Technical Writing SFT (100K)
100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use.
Motivation
Technical writing is one of the most underserved capabilities in LLMs. Common model failures:
Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.Align-Anything-Instruction-100K-zh
Dataset Card for Align-Anything-Instruction-100K-zh
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Instruction-Dataset-100K(zh)
Highlights
Data sources:
Firefly (47.8%),
COIG (2.9%),
and our meticulously constructed QA pairs (49.3%).
100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4o-mini.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
instruction-following-hard-sft-100k
Hard Instruction Following SFT (100K)
100,000 ShareGPT conversations where the assistant correctly satisfies multiple simultaneous explicit constraints in a single response. Each example pairs a multi-constraint prompt with a response that honors every constraint without dropping any.
Targets the instruction-following capability measured by IFEval and similar benchmarks.
Motivation
A key failure mode in deployed LLMs is dropping constraints under load — responding… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/instruction-following-hard-sft-100k.JumpLander-PMB-100K
🚀 JumpLander-PMB-100K
Persian Model Behavior Dataset for Intent, Constraint, and Safe Response Evaluation
جامپلندر PMB-100K | مجموعهداده فارسی برای سنجش رفتار مدل، فهم نیت، رعایت محدودیت و پاسخ امن
Built by JumpLander
Official Website: jumplander.orgPersian Website: jumplander.org/faDocumentation: jumplander.org/fa/docsAbout JumpLander: jumplander.org/fa/aboutSupport JumpLander: jumplander.org/fa/rateHugging Face Organization:… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpLander-PMB-100K.100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
cybersecurity-sft-100k
Cybersecurity SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering.
Dataset Description
This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.Align-Anything-Instruction-100K
Dataset Card for Align-Anything-Instruction-100K
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Highlights
Data sources:
PKU-SafeRLHF QA ,
DialogSum,
Empathetic,
Instruction-Wild,
and Alpaca.
100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.Chinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.honest-uncertainty-sft-100k
Honest Uncertainty SFT (100K)
100,000 ShareGPT conversations demonstrating calibrated epistemic humility across 21 scenarios. Each example shows a model correctly expressing what it knows, what it doesn't know, and why -- without being uselessly vague or confidently wrong.
Targets the hallucination and overconfidence failure modes that are the #1 complaint in enterprise AI deployments.
Motivation
LLMs have a systematic bias toward confident-sounding responses… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/honest-uncertainty-sft-100k.structured-output-sft-100k
Structured Output SFT (100K)
100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output.
Motivation
Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways:
Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.pii-masking-micro-100k
PII Masking Micro: Multilingual Sample
A micro-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.openpii-masking-micro-100k
OpenPII Micro: Multilingual PII Masking Sample
A micro-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
100,000
90,000
10,000
19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.Expert-Go-SFT-100K
Expert-Go-SFT-100K
Paper | Code
Expert-Go-SFT-100K is a large-scale synthetic dataset designed to "cold start" Large Language Models (LLMs) for Go-related reasoning tasks. It was introduced as part of the LoGos project, which aims to bridge the gap between general-purpose LLM reasoning and specialized expert knowledge in the game of Go.
The dataset features 100,000 samples of structured Go expertise mixed with general long Chain-of-Thought (CoT) reasoning data. It enables models to… See the full description on the dataset page: https://huggingface.co/datasets/YichuanMa/Expert-Go-SFT-100K.meeting-summarization-sft-100k
Meeting Summarization SFT (100K)
100,000 ShareGPT conversations demonstrating structured meeting summarization across 22 meeting types. Each example converts a realistic meeting transcript into a well-organized summary with key decisions, action items, and discussion notes — in the format that professional teams actually use.
Motivation
Meeting transcription tools (Otter.ai, Fireflies, Zoom AI) generate raw text but struggle to produce usable summaries. Common… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/meeting-summarization-sft-100k.vector-100k
VectorOS Vector 100k SimSat VLM Dataset
VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M.
The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.Muse-Glimmer-OPB-100K
Muse Glimmer OPB 100K
On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark.
Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B.
The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths.
Reasoning strength
Conversations
Train-turn rows
low
64,997
96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.devops-kubernetes-sft-100k
DevOps and Kubernetes SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality DevOps and Kubernetes conversations designed to train AI assistants capable of supporting platform engineers, SREs, and DevOps practitioners.
Dataset Description
This dataset covers production-grade Kubernetes operations, cloud infrastructure, CI/CD pipelines, GitOps workflows, and platform engineering across 13 specialized categories. Each record follows the ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/devops-kubernetes-sft-100k.pli-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
EPII Personal Location Information (PLI) Masking Preview Dataset
Overview
This dataset provides a preview (400 samples) of the EPII Personal Location Information (PLI) Masking Dataset, a… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pli-masking-100k.code-generation-sft-100k
Code Generation SFT (100K)
100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions.
Motivation
Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. This… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.pdi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
EPII Personal Digital Information (PDI) Masking Preview Dataset
Overview
This dataset provides a preview (400 samples) of the EPII Personal Digital Information (PDI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pdi-masking-100k.product-management-sft-100k
Product Management SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level product management across PRD writing, feature prioritization, OKR setting, roadmap planning, user research, competitive analysis, and stakeholder communication.
Motivation
AI assistants for product management commonly fail by:
Generic frameworks without application: Explaining RICE scoring without actually scoring the user's features; describing OKRs without writing them… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/product-management-sft-100k.pwi-masking-100k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
EPII Personal Work Information (PWI) Masking Preview Dataset
Overview
This dataset provides a preview (400 samples) of the EPII Personal Work Information (PWI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pwi-masking-100k.
