datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
math-reasoning-sft-100k
Math Reasoning SFT (100K)
100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models.
Dataset Description
100,000 problems across 8 mathematical categories and 3 difficulty levels:
Categories
Category
Examples
Topics
word_problems
~23,100
Rate/time/distance, work problems, mixture, meeting/catch-up
arithmetic
~15,400
Percentages, profit/loss, ratios
geometry
~15,400
Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.100k-corpus-2026
MAST 100K Corpus 2026
This dataset contains the fixed English document corpus used for MAST @ FIRE 2026, the Multilingual Agentic Search Track. MAST evaluates whether multilingual agentic search systems can answer complex questions posed in different languages by retrieving English evidence and producing short, correct English answers.
This corpus is copied from BrowseComp-Plus, a benchmark for Deep-Research systems that isolates the effect of the retriever and the LLM agent to… See the full description on the dataset page: https://huggingface.co/datasets/mast-benchmark/100k-corpus-2026.math-datasets-100k
Merged Math Datasets (100k Subset)
This dataset combines multiple mathematical datasets for training and evaluation purposes. This version contains a shuffled 100k subset of the training data for faster experimentation.
Dataset Description
A comprehensive collection of mathematical problems and solutions from various sources, organized into training and multiple test subsets.
Dataset Structure
Training Set
Size: 100000 examples
Fields: source… See the full description on the dataset page: https://huggingface.co/datasets/weijiezz/math-datasets-100k.medical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.science-qa-sft-100k
Science QA SFT (100K)
100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty.
Motivation
Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.MedDialog-EN-100kornstein-curated-100k
Ornstein Curated 100K
A curriculum-sorted reasoning dataset for SFT and post-training experiments.
Ornstein Curated 100K is a multi-domain instruction dataset built around explicit reasoning traces, difficulty progression, and curriculum-style ordering.
The dataset contains 100,000 samples across mathematics, programming, conversational reasoning, and cognitive-science-inspired tasks. It is sorted from easier to harder examples so users can train with the provided order, compare… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/ornstein-curated-100k.cybersecurity-sft-100k
Cybersecurity SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering.
Dataset Description
This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.Chinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.Sujet-Finance-QA-Vision-100k
Dataset Description 📊🔍
The Sujet-Finance-QA-Vision-100k is a comprehensive dataset containing over 100,000 question-answer pairs derived from more than 9,800 financial document images. This dataset is designed to support research and development in the field of financial document analysis and visual question answering.
Key Features:
🖼️ 9,801 unique financial document images
❓ 107,050 question-answer pairs
🇬🇧 English language
📄 Diverse financial document types… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-QA-Vision-100k.plot-palette-100k
Empowering Writers with a Universe of Ideas Plot Palette DataSet HuggingFace » Plot Palette was created to fine-tune large language models for creative writing, generating diverse outputs through iterative loops and seed data. It is designed to be run on a Linux system with systemctl for managing services. Included is the service structure, specific category prompts and ~100k data entries. The dataset is available here or… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/plot-palette-100k.devops-kubernetes-sft-100k
DevOps and Kubernetes SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality DevOps and Kubernetes conversations designed to train AI assistants capable of supporting platform engineers, SREs, and DevOps practitioners.
Dataset Description
This dataset covers production-grade Kubernetes operations, cloud infrastructure, CI/CD pipelines, GitOps workflows, and platform engineering across 13 specialized categories. Each record follows the ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/devops-kubernetes-sft-100k.AgentAngel_100k
Within Us AI — AgentAngel_100k (Agentic Coding 2026)
AgentAngel is a master-scholar, evidence-backed dataset family for training and evaluating agentic coding models that plan, patch, run checks, and iterate with tests-as-truth.
This release contains 100,000 examples per split (500,000 JSONL rows total):
Q&A (facts + rights/wrongs)
Instruct (messages)
Thinking (concise rationales)
Reasoning (constraints + verification checks)
Chat (multi-turn)
Evidence discipline… See the full description on the dataset page: https://huggingface.co/datasets/11-47/AgentAngel_100k.NuminaMath-100k
Merged Math Datasets (100k Subset)
This dataset combines multiple mathematical datasets for training and evaluation purposes. This version contains a shuffled 100k subset of the training data for faster experimentation.
Dataset Description
A comprehensive collection of mathematical problems and solutions from various sources, organized into training and multiple test subsets.
Dataset Structure
Training Set
Size: 100000 examples
Fields: source… See the full description on the dataset page: https://huggingface.co/datasets/weijiezz/NuminaMath-100k.browsecomp-plus-100k-corpus-as-local-folder
BrowseComp-Plus 100K Corpus — as local folder
The 100K-document subset of the BrowseComp-Plus
benchmark corpus, processed into a plain document-directory form
using the browsecomp-plus processing code from
DCI-Agent-Lite.
This repository stores the corpus exactly as it is expected on local disk:
a flat tree of <domain>/<title>.txt files, ready to be pointed at by
--corpus-dir. It is the form consumed by the RARG / DCI-Agent
retrieval-augmented agent during the BrowseComp-Plus… See the full description on the dataset page: https://huggingface.co/datasets/lossisnotanumber/browsecomp-plus-100k-corpus-as-local-folder.mlops-deployment-sft-100k
MLOps Deployment SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering MLOps and ML model deployment — from model serving and inference optimization to monitoring, CI/CD, and production scaling. Designed to train AI assistants that can help ML engineers deploy and operate models at scale.
Dataset Description
This dataset covers the full MLOps lifecycle across 13 specialized categories. Each record follows the ShareGPT… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/mlops-deployment-sft-100k.DistilQwen_100k_korean
DistilQwen 100k Korean
This dataset is a Korean translation of the original alibaba-pai/DistilQwen_100k dataset.
Dataset Structure
The dataset contains both English and Korean versions of instruction-response pairs:
{
"instruction": "Original English instruction text",
"output": "Original English response/answer",
"instruction_kr": "Korean translation of the instruction",
"output_kr": "Korean translation of the response/answer",
"_dataset_index": 30000
}
Each… See the full description on the dataset page: https://huggingface.co/datasets/lcw99/DistilQwen_100k_korean.bitcoin-security-reasoning-100k
Dataset Card for Bitcoin Security Reasoning 100K
100,000 high-quality synthetic training samples for fine-tuning LLMs on Bitcoin protocol security analysis. Teaches models to analyze vulnerability clusters, form security hypotheses, and generate differential testing code.
Dataset Details
Dataset Description
This dataset contains structured security reasoning chains for Bitcoin protocol vulnerabilities. Each sample presents a cluster of causal… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/bitcoin-security-reasoning-100k.sefd-archive-100k-analysis-sample-qwen3-20260524
SEFD Archive 100k Analysis Sample Qwen3 20260524
Retained artifacts for the completed archive-wide 100,000-filing Stanford EDGAR Filings Dataset (SEFD) analysis sample used in the arXiv paper update. The sample contains 2,971,490,909 final SEFD tokens, counted with the Qwen3-1.7B tokenizer.
This repository is a new versioned artifact and intentionally does not replace the earlier sfd-archive-100k-analysis-sample repository used for the original conference submission.
Included:… See the full description on the dataset page: https://huggingface.co/datasets/sfd-anonymous/sefd-archive-100k-analysis-sample-qwen3-20260524.wikipedia-qa-ja-100k
Dataset Card for "wikipedia-qa-ja-100k"
Original Dataset
hpprc/wikipedia-20240101
Procedure
Extract the first line of the title from the dataset.
Generate the answer by summizing the line using LLM:
Input RAG-like prompt to CALM 2 7B Chat.
Format the response.
RAG-like Prompt
f"""USER: {title}とはなんですか?次の文章を参考に一言でまとめてください。{text}
ASSISTANT: """
rlhn-100K
Dataset Card for RLHN-100K
Dataset Description
Repository |
Paper |
ArXiv
RLHN is a cascading LLM framework designed to accurately relabel hard negatives in existing IR/RAG training datasets, such as MS MARCO and HotpotQA.
This Tevatron dataset (100K training pairs) contains the queries, positives + relabeled hard negatives, remaining hard negatives for 7 datasets in the BGE training collection.
This repository contains the training pairs that can be used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/rlhn/rlhn-100K.Genesis_AI_Code_100k
Genesis AI Code 100K (Frontier)
Developed by: Within Us AI
Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation.
Splits
train: 98,000
validation: 2,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_100k.100k_Tdk_zurriyet_dna_v6.jsonl
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/100k_Tdk_zurriyet_dna_v6.jsonl.Genesis_AI_Code_100k
Genesis AI Code 100K (Frontier)
Developed by: Within Us AI
Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation.
Splits
train: 98,000
validation: 2,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Genesis_AI_Code_100k.HealthTalks-100kslimorca-dedup-chatml-100k
Copy of Open-Orca/SlimOrca-Dedup in ChatML format downsample to 100k
"SlimOrca Dedup" is a deduplicated, unfiltered subset of the SlimOrca dataset, excluding RLHF instances, resulting in 363k unique examples.
Key Features
Removal of RLHF instances.
Deduplication using minhash and Jaccard similarity techniques.
Demo Models
Note: These models were trained on the full SlimOrca dataset, not the deduplicated, unfiltered version.
*… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/slimorca-dedup-chatml-100k.rag-systems-sft-100k
RAG Systems SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering Retrieval-Augmented Generation (RAG) systems — from basic pipelines to advanced multi-hop retrieval, evaluation, and production optimization. Designed to train AI assistants that can help engineers build, debug, and scale RAG applications.
Dataset Description
This dataset covers the full spectrum of RAG system development across 12 specialized categories.… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-systems-sft-100k.hn-remove-100K
Dataset Card for HN-Remove 100K
Dataset Description
Repository |
Paper |
ArXiv
RLHN is a cascading LLM framework designed to accurately relabel hard negatives in existing IR/RAG training datasets, such as MS MARCO and HotpotQA.
This Tevatron dataset (100K training pairs) contains the queries, positives, hard negatives (with dropped false negatives) for 7 datasets in the BGE training collection.
This repository contains the training pairs that can be used to fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/rlhn/hn-remove-100K.Reasoning-Mix-100k
🧠 Reasoning Mix 100k ✨
A high-quality, balanced reasoning dataset consisting of 99,999 samples extracted from three reasoning datasets. This dataset is specifically formatted for models to utilize a /think block for step-by-step reasoning.
📊 Dataset Summary
The Reasoning Mix 100k is a curated collection of reasoning tasks, primarily focused on mathematics, logic, and general problem-solving. It combines the strengths of three high-performing datasets into a unified… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Reasoning-Mix-100k.Pashto-100k-Pairs
Qehwa AI - Pashto 100K Fine-Tuning Dataset
Overview
Qehwa AI presents a large-scale Pashto instruction tuning dataset containing 100,000+ high-quality instruction-response pairs designed for supervised fine-tuning, conversational AI, and downstream NLP tasks.
This dataset was created to advance AI research for the Pashto language, a significantly underrepresented low-resource language spoken by millions worldwide. The dataset covers more than 20 diverse domains and is… See the full description on the dataset page: https://huggingface.co/datasets/junaid008/Pashto-100k-Pairs.
