datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ui-instruct-4k
UI Instruct 4K
A instruction-completion dataset for finetuning language models to specialize in generating Next.js / ShadCN UI components using React, TypeScript, and Tailwind CSS.
Dataset Summary
This dataset was created with the primary goal of finetuning Qwen 3.5 4B to become a specialist at outputting production-ready Next.js and ShadCN-based UI components. Each example consists of a natural language prompt describing a UI component or layout, paired with a clean… See the full description on the dataset page: https://huggingface.co/datasets/iamdyeus/ui-instruct-4k.marketing-instruct-4k
Dataset Card for marketing-instruct-4k
Dataset Details
Dataset Description
A curated instruction-tuning dataset of ~4,300 marketing copywriting
examples across five task types, built for the AutoScientist Challenge
2026 (Marketing category). Used to fine-tune Marketing-Mixtral-8x7B.
Key finding: this carefully curated dataset at its natural size
outperformed a 12,000-row version expanded via automated augmentation
(80% vs 58% win rate against the… See the full description on the dataset page: https://huggingface.co/datasets/suehuynh/marketing-instruct-4k.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structured-hard-sft-4k.Multilingual-FanFic-Chat-4K
Dataset Card for Multilingual FanFic Chat 4K
A dataset of 4,000 chat interactions for fan fiction in multiple languages.
Dataset Details
Dataset Description
This dataset consists of 4,000 simulated chat interactions specifically designed to assist in writing fan fiction in 40 different languages. The interactions were generated using GPT-3.5 Turbo and include both questions and responses related to fan fiction writing for various popular properties.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/robgonsalves/Multilingual-FanFic-Chat-4K.Organized_4k_Contex_WIUAI_1.3M
ITLL/Organized_PreTrain_WIUAI
A clean, deduped, 4096-context pretrain-ready merge of 31 distillation datasets from 11-47 and WithinUsAI.
All examples >4096 tokens were 100% trashed, never truncated. Duplicates were removed globally across all sources. Outputs are grouped by category for curriculum pretraining.
Total Kept: 1,372,185 examples
Total Dropped >4096: 1,208 examples
Total Duplicates Dropped: ~181k+ (including shard/monolith duplicates)
Context: ≤4096 tokens… See the full description on the dataset page: https://huggingface.co/datasets/ITLL/Organized_4k_Contex_WIUAI_1.3M.msm-packaging-claude-green-chatgpt-blue-4k5-v3
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-4k5-v3
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.structured-hard-sft-4k
Hard Synthetic Dataset for Structured Data Tasks (v1)
This dataset contains 4,000 high-difficulty synthetic samples designed to improve LLM performance on complex structured data conversion, extraction, and formatting tasks.
The data is fully synthetic, generated using deterministic serialization to ensure syntax validity while maintaining high structural complexity (deep nesting and varied types).
Dataset Summary
The dataset addresses four "hard" areas typically… See the full description on the dataset page: https://huggingface.co/datasets/zyz123code/structured-hard-sft-4k.smolmo-olmo3-calib-4k
Olmo-3 PTQ Calibration Set (4k, math)
A 4,000-sample calibration set for post-training quantization (FP8 / NVFP4) of
allenai/Olmo-3.1-32B-Think. Each row is a
complete math reasoning conversation rendered with the Olmo-3 chat template (the text field),
so calibration sees exactly the model's native inference format — <|im_start|> turn markers, the
Olmo system prompt, <think>…</think> traces, and tool-use scaffolding.
How it was built
Sampled 4 random examples… See the full description on the dataset page: https://huggingface.co/datasets/chankhavu/smolmo-olmo3-calib-4k.msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B
The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both
name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B
The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.hallucination-grounding-dpo-4k
Hallucination Grounding DPO Pairs (4K)
DPO preference pairs targeting the full spectrum of factuality failures — from hallucination to over-hedging.
Motivation
Existing refusal/safety datasets focus on what not to say. This dataset targets the orthogonal challenge: when to say "I don't know" vs. when to answer confidently. Models that over-refuse waste user trust; models that hallucinate destroy it.
Dataset Description
4,000 preference pairs across… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/hallucination-grounding-dpo-4k.ipc_decisions_4kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
thwiki-2026-super-clean-4k
🇹🇭 Thai Wikipedia Super-Clean (2026 Edition)
Dataset ชุดนี้สกัดจาก Wikipedia ภาษาไทย (Dump 2026) โดยเน้นความสะอาดระดับ "Pro-Clean" เพื่อใช้สำหรับ Distillation และ Fine-tuning LLM โดยเฉพาะ
Key Features
High Information Density: คัดเฉพาะบทความที่มีเนื้อหายาวเกิน 1,000 ตัวอักษร และมีสัดส่วนภาษาไทย > 50%
Pro-Cleaned: ลบชื่อไฟล์ภาพ (.jpg, .png), ขยะสัญลักษณ์ Wiki (==, '''), และวงเล็บเปล่าออกทั้งหมด
Entity Preserved: เก็บเครื่องหมายคำพูด " "… See the full description on the dataset page: https://huggingface.co/datasets/bombman/thwiki-2026-super-clean-4k.agentic-llm-pretraining-1.7b-tokenized-qwen3-4k
Agentic LLM Pretraining Dataset - Tokenized (Qwen3, 4K context)
Pre-tokenized version of visionscaper/agentic-llm-pretraining-1.7b for pre-training small language models for agentic AI use cases.
Overview
Property
Value
Source dataset
visionscaper/agentic-llm-pretraining-1.7b
Tokenizer
Qwen/Qwen3-1.7B
Context length
4,096 tokens
EOD token
<|endoftext|> (ID 151643)
Token dtype
uint32
Total samples
375,384
Total tokens
~1.54 billion
Storage
~5.8 GB… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b-tokenized-qwen3-4k.ipc_decisions_4k_1024Датасет судебных решений суда по интеллектуальным правам РФ со строками до 1024 символов и синтаксисом для дообучения с инструкциями.
PyMath-AgentRL-4K
🧮 PyMath-AgentRL-4K
A supervised fine-tuning (SFT) dataset for training language models to solve mathematical problems using Python code interpreters.
Overview
Total Examples: 4,000 (deduplicated)
Format: Parquet
Language: English
License: Apache 2.0
Source Datasets
Merged from two high-quality sources:
Gen-Verse/Open-AgentRL-SFT-3K (3,000 examples)
JoeYing/ReTool-SFT (2,000 examples)
Deduplication: 20% overlap removed (1,000 duplicates)
Key… See the full description on the dataset page: https://huggingface.co/datasets/SimonSun/PyMath-AgentRL-4K.reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs.
The dataset has these columns for users to filter out:
repo_id
tok_len
user
thought_trace
assistant
ChatML
Repositories… See the full description on the dataset page: https://huggingface.co/datasets/willworker/reasoning-corpus-4K-5M-v1.reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs.
The dataset has these columns for users to filter out:
repo_id
tok_len
user
thought_trace
assistant
ChatML
Repositories… See the full description on the dataset page: https://huggingface.co/datasets/danie1111/reasoning-corpus-4K-5M-v1.Code-Reasoning-4k
Code-Reasoning-4k
A curated dataset for code-centric reasoning tasks, combining programming problems, mathematical reasoning, and general instruction-following samples. Despite the name, the dataset currently contains 38,140 samples, reflecting a significant expansion beyond the original “4k” scale.
Overview
This dataset is designed to support training and evaluation of models on:
Code generation and debugging
Algorithmic reasoning
Mathematical problem solving
Light… See the full description on the dataset page: https://huggingface.co/datasets/nphearum/Code-Reasoning-4k.ipc_decisions_4k_selectedДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
toolrl-4k-verl
ToolRL Dataset - GPT OSS 120B Format
A preprocessed tool-learning dataset in GPT OSS 120B native format for reinforcement learning training with GRPO/PPO algorithms.
Dataset Description
This dataset contains 4,000 tool-use samples (3,920 training / 80 test) converted from the ToolRL dataset to GPT OSS 120B's native format. The conversion replaces XML-style tags with GPT OSS's special tokens and channel system, resulting in ~10-15% token efficiency improvement.… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/toolrl-4k-verl.rag-grounding-dpo-4k
RAG Grounding DPO Pairs (4K)
DPO preference pairs for training LLMs to faithfully use retrieved context in RAG pipelines.
Motivation
RAG is the dominant LLM deployment pattern in production. The core failure mode: models that ignore retrieved context and hallucinate answers, or that contradict documents with confidently-stated fabrications. This dataset trains models to ground answers in provided context.
Dataset Description
4,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/rag-grounding-dpo-4k.ipc_decisions_4k_2048Датасет судебных решений суда по интеллектуальным правам РФ со строками до 2048 символов и синтаксисом для дообучения с инструкциями.
reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs.
The dataset has these columns for users to filter out:
repo_id
tok_len
user
thought_trace
assistant
ChatML
Repositories… See the full description on the dataset page: https://huggingface.co/datasets/tonySlowWriter/reasoning-corpus-4K-5M-v1.Qyrou-reasoning-corpus-4K-5M-v1 Reasoning Corpus 5M · Within 5k sequence length
About Dataset
This dataset contains reasoning chains from major AI models, such as: DeepSeek-v4 (both Pro and Flash), DeepSeek-r1 (DS-r1, Llama-DS, Qwen-DS), Qwen3, Qwen3.5/3.6 (both OpenSource and API models), Gemma4-31B derived from many other repositories, and properly filtered to train SLMs.
The dataset has these columns for users to filter out:
repo_id
tok_len
user
thought_trace
assistant
ChatML
Repositories… See the full description on the dataset page: https://huggingface.co/datasets/vietdata/Qyrou-reasoning-corpus-4K-5M-v1.Databricks-Dolly-4k
Databricks-Dolly-4k
The resulting dataset contains 4000 samples of the databricks/databricks-dolly-15k dataset.
This split of an even smaller subset is provided for very fast experimentation and evaluation of models when computational resources are highly limited or for quick prototyping.
Dataset Structure
The dataset is provided as a DatasetDict with the following splits:
train: Contains 4000 samples.
Each split contains the following features, identical to the… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Databricks-Dolly-4k.gpt2-general-qa-4kThis is a general Dataset for basic GPT2 fine tuning, with instructions and answers.
