datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLaVA-CoT-100k
Dataset Card for LLaVA-CoT
The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.HealthCareMagic-100k-enOmniVideo-100K
OmniVideo-100K
Official repository for OmniVideo-100K, an instruction-tuning dataset introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos.tar.part_xx: Raw video files.
train_oe_70k.jsonl: Original Open-Ended (OE) training samples.
train_mcq_30k.jsonl: Original Multiple-Choice (MCQ) training samples.
train_oe_70k_formatted.jsonl: Instruction-formatted OE samples (ready for… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-100K.math-reasoning-sft-100k
Math Reasoning SFT (100K)
100,000 math problems with detailed step-by-step solutions — ready for supervised fine-tuning of math reasoning models.
Dataset Description
100,000 problems across 8 mathematical categories and 3 difficulty levels:
Categories
Category
Examples
Topics
word_problems
~23,100
Rate/time/distance, work problems, mixture, meeting/catch-up
arithmetic
~15,400
Percentages, profit/loss, ratios
geometry
~15,400
Area… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/math-reasoning-sft-100k.UniREdit-Data-100KUniREditBench: A Unified Reasoning-based Image Editing Benchmark
medical-instruction-100k
What is the Dataset About?🤷🏼♂️
The dataset is useful for training a Generative Language Model for the Medical application and instruction purposes, the dataset consists of various thoughs proposed by the people [mentioned as the Human ] and there responses including Medical Terminologies not limited to but including names of the drugs, prescriptions, yogic exercise suggessions, breathing exercise suggessions and few natural home made prescriptions.
How the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed-Altaf/medical-instruction-100k.tool-calling-english-100k
Tool Calling English (100K)
100,000 tool-calling conversations in OpenAI function calling format — the largest general English tool-use dataset for fine-tuning.
Motivation
Models trained without tool-calling examples struggle in agentic deployments. This dataset trains the full cycle: deciding when to call a tool, calling it with correct arguments, interpreting the result, and producing a grounded final response.
Dataset Description
100,000… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/tool-calling-english-100k.DCLM-200-100k-exact-dedupExpert-Sudoku-100kmedical-clinical-reasoning-sft-100k
Medical Clinical Reasoning SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality medical and clinical reasoning conversations designed to train AI assistants capable of supporting clinical decision-making, documentation, and medical education.
Dataset Description
This dataset covers a broad spectrum of clinical practice scenarios across 10 medical specialty categories. Each record follows the ShareGPT conversation format with a detailed human… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/medical-clinical-reasoning-sft-100k.VideoInstruct-100KVideoInstruct100K is a high-quality video conversation dataset generated using human-assisted and semi-automatic annotation techniques. The question answers in the dataset are related to,
Video Summariazation
Description-based question-answers (exploring spatial, temporal, relationships, and reasoning concepts)
Creative/generative question-answers
For mored details, please visit Oryx/VideoChatGPT/video-instruction-data-generation.
If you find this dataset useful, please consider citing the… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/VideoInstruct-100K.HealthCareMagic-100k-Chat-Format-enwriting-quality-dpo-100k
Writing Quality DPO (100K)
100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect.
Motivation
Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.recommendations-ml-100k
MovieLens Leave-One-Out
Five chronological interactions predict the next interaction. One final test target per user; no rating filter. Histories are audit-only, not wholesale model inputs. Actors are supplementary; see actor_sources.json.
{
"schema": "movie-fields-v1",
"source": "official MovieLens 100K",
"sample_policy": "leave-one-out-windows",
"past_order": "oldest-first",
"history_length": 5,
"stride": 1,
"shuffle_seed": 42,
"timestamp_policy": "rating… See the full description on the dataset page: https://huggingface.co/datasets/Nithish2410/recommendations-ml-100k.email-writing-sft-100k
Email Writing SFT (100K)
100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails.
Motivation
Email is the primary communication channel for most professional work, yet LLMs often produce emails that are:
Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.science-qa-sft-100k
Science QA SFT (100K)
100,000 science Q&A examples with step-by-step explanations for SFT fine-tuning. Covers physics, chemistry, biology, astronomy, and earth science at beginner through advanced difficulty.
Motivation
Models trained on general text often give superficially plausible but mechanistically wrong answers to science questions — stating the right conclusion without understanding the underlying reasoning. This dataset trains models to explain why an… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/science-qa-sft-100k.lavita-ChatDoctor-HealthCareMagic-100ksn38-quality-gold-100k
SN38 quality prompt + gold continuation
Synthetic incomplete-sentence prompts with gold continuations for Bittensor
subnet 38. 13 categories, 13 items per category per call,
temperature 1.0.
Each row:
category: reading_comprehension, language_understanding, world_knowledge, commonsense_reasoning, language_modeling, causal_reasoning, logical_inference, temporal_reasoning, math_reasoning, truthfulness, pronoun_resolution, paraphrase_detection, word_sense_disambiguation
prompt:… See the full description on the dataset page: https://huggingface.co/datasets/jjjlimaus/sn38-quality-gold-100k.technical-writing-sft-100k
Technical Writing SFT (100K)
100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use.
Motivation
Technical writing is one of the most underserved capabilities in LLMs. Common model failures:
Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.IMIG-100K
IMIG-100K: A Large-Scale Synthetic Dataset for Multi-Instance Image Generation with Detailed Annotation
Ruihang Xu,
Dewei Zhou,
Fan Ma†,
Yi Yang
ReLER Lab, CCAI, Zhejiang University
📄 Dataset Overview
The IMIG-100K dataset is a large-scale synthetic dataset designed for multi-instance image generation tasks. It contains more than 100,000 high-quality image samples, each annotated with masks and layout information. The dataset is organized into several sub-datasets… See the full description on the dataset page: https://huggingface.co/datasets/ruihangxu/IMIG-100K.Align-Anything-Instruction-100K-zh
Dataset Card for Align-Anything-Instruction-100K-zh
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Instruction-Dataset-100K(zh)
Highlights
Data sources:
Firefly (47.8%),
COIG (2.9%),
and our meticulously constructed QA pairs (49.3%).
100K QA pairs (zh): 104,550 meticulously crafted instructions, selected and polished from various Chinese datasets… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K-zh.100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4o-mini.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
instruction-following-hard-sft-100k
Hard Instruction Following SFT (100K)
100,000 ShareGPT conversations where the assistant correctly satisfies multiple simultaneous explicit constraints in a single response. Each example pairs a multi-constraint prompt with a response that honors every constraint without dropping any.
Targets the instruction-following capability measured by IFEval and similar benchmarks.
Motivation
A key failure mode in deployed LLMs is dropping constraints under load — responding… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/instruction-following-hard-sft-100k.JumpLander-PMB-100K
🚀 JumpLander-PMB-100K
Persian Model Behavior Dataset for Intent, Constraint, and Safe Response Evaluation
جامپلندر PMB-100K | مجموعهداده فارسی برای سنجش رفتار مدل، فهم نیت، رعایت محدودیت و پاسخ امن
Built by JumpLander
Official Website: jumplander.orgPersian Website: jumplander.org/faDocumentation: jumplander.org/fa/docsAbout JumpLander: jumplander.org/fa/aboutSupport JumpLander: jumplander.org/fa/rateHugging Face Organization:… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpLander-PMB-100K.100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds.
We generate the dataset with the following steps:
Generate 120k descriptions by GPT4.
Generate 120k codes follow each description by GPT4o-mini.
Run the 120k codes and do auto-filtering.
Get the final 100k legitimate ARC-like tasks with examples.
cybersecurity-sft-100k
Cybersecurity SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering.
Dataset Description
This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.Align-Anything-Instruction-100K
Dataset Card for Align-Anything-Instruction-100K
[🏠 Homepage]
[🤗 Instruction-Dataset-100K(en)]
[🤗 Instruction-Dataset-100K(zh)]
[🤗 Align-Anything Datasets]
Highlights
Data sources:
PKU-SafeRLHF QA ,
DialogSum,
Empathetic,
Instruction-Wild,
and Alpaca.
100K QA pairs: By leveraging GPT-4 to annotate meticulously refined instructions, we obtain 105,333 QA pairs.… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Align-Anything-Instruction-100K.MedDialog-EN-100kChinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.honest-uncertainty-sft-100k
Honest Uncertainty SFT (100K)
100,000 ShareGPT conversations demonstrating calibrated epistemic humility across 21 scenarios. Each example shows a model correctly expressing what it knows, what it doesn't know, and why -- without being uselessly vague or confidently wrong.
Targets the hallucination and overconfidence failure modes that are the #1 complaint in enterprise AI deployments.
Motivation
LLMs have a systematic bias toward confident-sounding responses… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/honest-uncertainty-sft-100k.
