datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
creative-writing-sft-50k
Creative Writing SFT (50K)
50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure.
Motivation
Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed.
(Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined.
Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.RL-Claude-Creative-Writing-SFT
RL-Claude-Creative-Writing-SFT
Alpaca-format dataset. Columns: instruction, input, output
from datasets import load_dataset
ds = load_dataset("SLoonker/RL-Claude-Creative-Writing-SFT", split="train")
wildchat_creative_writing_annotated_10kCreative-Writing-High-Quality-1300x
Creative Writing - Part One (Shadow & Skeleton)
This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology.
Methodology: Shadow & Skeleton
Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach:
Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course).
This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.smoltalk-creative-writingCreative-Writing-Gemini3Pro-2700x
Pulitzer Diamond Prose GEMINI Seeds
This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.Japanese-Creative-Writing-39.6k
Japanese-Creative-Writing-39.6k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約39600件の日本語の小説執筆タスクデータセットです。
全てのデータは2ターンのデータとなっています。また、データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction_1: 1ターン目の指示プロンプト
output_1: 1ターン目のアシスタント応答
instruction_2: 2ターン目の指示プロンプト
output_2: 2ターン目のアシスタント応答
1ターン目の指示プロンプトはdeepseek-ai/DeepSeek-V3-0324で合成されています。system promptや2ターン目の指示プロンプトは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-Creative-Writing-39.6k.Creative-Writing-Sonnet4.6-800x
Pulitzer Diamond Prose CLAUDE Seeds
This dataset contains 833 high-quality creative writing seeds generated using Claude 4.6 Sonnet.
Each entry represents a story opening designed to meet high literary standards.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements: extreme show-don't-tell, double-labor sentence… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-800x.RL-Claude-Creative-Writing-DPO
RL-Claude-Creative-Writing-DPO
Alpaca-format dataset. Columns: instruction, input, output, rejected
from datasets import load_dataset
ds = load_dataset("SLoonker/RL-Claude-Creative-Writing-DPO", split="train")
Creative-Writing-Sonnet4.6-Cleaned
Creative-Writing-Sonnet4.6-Cleaned
Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved.
Format
Each line is a JSON object with:
messages: list of message dicts with roles (system, user, assistant)
System: writing quality instructions
User: cleaned creative writing prompt
Assistant: creative writing response (may include <think> traces)
Stats
Metric
Value
Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.Creative-Writing-Part-Two
Creative Writing - Part Two (The Nuclear Dataset)
This dataset represents the "Nuclear" layer of our creative writing training pipeline. While Part One focused on physical and psychological grounding (Shadow & Skeleton), Part Two focuses on dense literary resonance, subtext, and stylistic sophistication.
Methodology: The Nuclear Pipeline
This dataset was built using a multi-phase "Controlled Criticality" approach to ensure maximum signal density without the… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Part-Two.dolly_creative_writing
Dataset Card for "dolly_creative_writing"
More Information needed
wildbench-creative-writingCreative-Writing-KimiK2.5-Cleaned
Creative-Writing-KimiK2.5-Cleaned
Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved.
Format
Each line is a JSON object with:
messages: list of message dicts with roles (system, user, assistant)
System: writing quality instructions
User: cleaned creative writing prompt
Assistant: creative writing response (may include <think> traces)
Stats
Metric
Value
Total prompt tokens
80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.Creative-Writing-Qwen3.5Plus-2000x
Pulitzer Diamond Prose QWEN Seeds
This dataset contains 2638 high-quality creative writing seeds generated using Qwen 2.5 72B.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Qwen3.5Plus-2000x.Creative-Writing-Multiturn-Cleaned-16kOrion-Creative_Writing-Complexitycqa-creative-writing-expert-cot-preview
CQA: Creative Quality Alignment — Research-Grade Schema v2
English
This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.creative_writing_conversationwildchat-creative-writing-3k-rftcreative_writing
Dataset Card for telecomadm1145/creative_writing
Dataset Details
Dataset Description
This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output).
The dataset emphasizes:
Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.creative_writingcreative_writing_v9Creative-Writingwildchat-creative-writing-3k-prefCreative_Writing_ShareGPT_Enhanced
Creative Writing ShareGPT — Enhanced Edition ✨
High-quality creative writing dataset with regenerated responses using StepFun's Step-3.5-Flash model.
This dataset is an enhanced version of ChaoticNeutrals/Creative_Writing-ShareGPT, where all final AI responses have been regenerated using stepfun/step-3.5-flash with a carefully engineered system prompt designed to produce literary-quality creative writing.
What Changed
Original human prompts preserved — All… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative_Writing_ShareGPT_Enhanced.Creative-Writing-Reasoning-KimiK2.5-600x
Pulitzer Diamond Prose KIMI Seeds
This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.eq-bench-creative-writing-v3https://eqbench.com/creative_writing.html
eq-bench
