datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WRIT-2K
WRIT-2K
WRIT-2K is a 2,000-trajectory supervised fine-tuning dataset for multi-turn, tool-using customer-service agents on tau2-bench style airline and retail tasks.
This dataset accompanies the paper WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents.
Project homepage: https://hengrui-gu.github.io/WRIT/
Dataset Summary
WRIT-2K contains complete multi-turn trajectories with user messages, assistant natural-language responses… See the full description on the dataset page: https://huggingface.co/datasets/Henryoung/WRIT-2K.HQ-Chat-2k
🧠 HQ-Chat-2K — High-Quality Conversational & Instruction-Tuning Dataset
2,000 carefully curated, high-quality conversation and instruction examples for fine-tuning Small Language Models (SLMs) and compact LLMs from ~500M to 3B parameters.
HQ-Chat-2K is a high-quality conversational and instruction-tuning dataset designed specifically for training and fine-tuning small to medium-sized Large Language Models (LLMs).
The dataset contains 2,000 curated user–assistant examples… See the full description on the dataset page: https://huggingface.co/datasets/ThinkNet/HQ-Chat-2k.agentic-foresight-actions-2k
Agentic Foresight: 2K Multi-Step JSON Action & Rollback Dataset
Dataset Description
This dataset contains 2,000 highly structured, synthetically generated input/output pairs explicitly designed to train Large Language Models in Agentic Foresight, Multi-Step Orchestration, and Sequential Task Automation.
Unlike standard tool-calling datasets that map a single prompt to a single API call, this dataset forces the model to act as a macro-orchestrator. It translates… See the full description on the dataset page: https://huggingface.co/datasets/Qapdex/agentic-foresight-actions-2k.cot-reasoning-2k
DuoNeural CoT Reasoning Dataset (2K)
A compact, high-quality chain-of-thought reasoning dataset generated for supervised fine-tuning (SFT). All 2,151 examples are quality-scored 5/5 and focus on explicit step-by-step reasoning traces.
Benchmark Results
Fine-tuned Qwen2.5-1.5B-Instruct on this dataset (3 epochs, LoRA rank 16, ~36 min on RTX 3090):
Metric
Baseline
Post-SFT
Δ Absolute
Δ Relative
GSM8K (flexible-extract)
0.3177
0.4890
+17.1pp
+53.9%
GSM8K… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/cot-reasoning-2k.gpt-5.4-xhigh-reasoning-2k
Gpt-5.4-Xhigh-Reasoning-2000x
A premium-quality reasoning dataset containing 2,007 elite samples distilled from GPT-5.4 XHIGH (the highest reasoning effort tier of GPT-5.4). Each sample features deep, multi-step Chain-of-Thought traces that are significantly longer and more rigorous than standard GPT-5.4 outputs.
This dataset is specifically designed for Supervised Fine-Tuning (SFT) to transform general-purpose language models into powerful reasoning models with explicit… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5.4-xhigh-reasoning-2k.chain-of-thought-dpo-2k
Chain-of-Thought DPO Pairs (2.6K)
DPO preference pairs for training LLMs to reason explicitly before answering.
Dataset Description
2,600 preference pairs across 6 reasoning categories:
Category
Examples
Description
math_word
~610
Multi-step math word problems
coding
~420
Algorithm complexity, CS reasoning
economics
~415
Economic analysis and theory
science
~390
Physics, chemistry, biology reasoning
logic
~390
Deductive reasoning, puzzles… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/chain-of-thought-dpo-2k.ro-preference-pairs-2k
Romanian Preference Pairs (2K)
Synthetic DPO preference pairs in Romanian targeting over-refusal and helpfulness alignment.
Dataset Description
2,000 preference pairs in Romanian across 4 categories:
Category
Examples
Description
informational
~500
Factual questions about Romania, economics, law
coding
~500
Python code tasks, FastAPI, SQLAlchemy
task_completion
~500
Document drafting, emails, plans
advice
~500
Career, productivity, technical… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/ro-preference-pairs-2k.deepseek_cot_2k
Deepseek CoT 2k
This dataset contains 1,515 extracted records focused on Chain-of-Thought (CoT) reasoning. It was processed from a malformed JSON source and converted into a clean, ready-to-use JSONL format.
Dataset Structure
Each record follows this schema:
id: Unique identifier for the sample.
problem: The input prompt or question.
thinking: The internal reasoning or "Chain of Thought" process.
solution: The final concise answer.
difficulty: Categorization of… See the full description on the dataset page: https://huggingface.co/datasets/3amthoughts/deepseek_cot_2k.thwiki-2026-super-clean-2k
🇹🇭 Thai Wikipedia Super-Clean (2026 Edition)
Dataset ชุดนี้สกัดจาก Wikipedia ภาษาไทย (Dump 2026) โดยเน้นความสะอาดระดับ "Pro-Clean" เพื่อใช้สำหรับ Distillation และ Fine-tuning LLM โดยเฉพาะ
Key Features
High Information Density: คัดเฉพาะบทความที่มีเนื้อหายาวเกิน 1,000 ตัวอักษร และมีสัดส่วนภาษาไทย > 50%
Pro-Cleaned: ลบชื่อไฟล์ภาพ (.jpg, .png), ขยะสัญลักษณ์ Wiki (==, '''), และวงเล็บเปล่าออกทั้งหมด
Entity Preserved: เก็บเครื่องหมายคำพูด " "… See the full description on the dataset page: https://huggingface.co/datasets/bombman/thwiki-2026-super-clean-2k.Somali-Somlish-Instruct-2K-Dataset
Somlish-Tech-Instruct-2K
This is the first-of-its-kind Somlish (Somali + English) instruction-tuning dataset. It contains 2,312 rows of high-quality synthetic data generated to teach AI models how to speak like a modern Somali tech enthusiast.
🌟 Why this exists
Standard Somali datasets are often too formal. This dataset uses natural "Discord-style" slang (Niyo, Sxb, Bro) while maintaining English technical terms (API, GPU, React) to ensure the AI stays smart and logical.… See the full description on the dataset page: https://huggingface.co/datasets/Zyroxx66/Somali-Somlish-Instruct-2K-Dataset.SciDataSailor-SFT-2K
SciDataSailor-SFT-2K
A scientific data exploration SFT dataset in JSONL format.
Files
SciDataSailor-SFT-2K.jsonl: SFT training samples.
Format
Each line is a JSON object.
SciDataSailor-SFT-2K
SciDataSailor-SFT-2K
A scientific data exploration SFT dataset in JSONL format.
Files
SciDataSailor-SFT-2K.jsonl: SFT training samples.
Format
Each line is a JSON object.
