datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
creative-writing-sft-50k
Creative Writing SFT (50K)
50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure.
Motivation
Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.CommonCrawl-CreativeCommons
The Common Crawl Creative Commons Corpus (C5)
Raw CommonCrawl crawls, annotated with Creative Commons license information
C5 is an effort to collect Creative Commons-licensed web data in one place.
The licensing information is extracted from the web pages based on whether they link to Creative Commons licenses either overtly in a tags (like in the footer of Wikipedia) or in metadata fields indicating deliberate Creative Commons publication. However, false positives may occur! See… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons.Creative-Writing-High-Quality-1300x
Creative Writing - Part One (Shadow & Skeleton)
This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology.
Methodology: Shadow & Skeleton
Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach:
Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.Creative-Writing-Gemini3Pro-2700x
Pulitzer Diamond Prose GEMINI Seeds
This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.CommonCrawl-CreativeCommons-fine
Common Crawl Creative Commons Corpus Fine (C5f)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that are also present in the FineWeb or FineWeb-2 datasets. As such, this dataset contains a high-quality subset of C5.
Created with this script.
For more information, see C5.
Progress
In the v1 release, the following crawls are included
CC-MAIN-2019-30
CC-MAIN-2020-05CC-MAIN-2023-06
CC-MAIN-2024-51
CC-MAIN-2024-46
CC-MAIN-2025-05… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-fine.Japanese-Creative-Writing-39.6k
Japanese-Creative-Writing-39.6k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約39600件の日本語の小説執筆タスクデータセットです。
全てのデータは2ターンのデータとなっています。また、データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction_1: 1ターン目の指示プロンプト
output_1: 1ターン目のアシスタント応答
instruction_2: 2ターン目の指示プロンプト
output_2: 2ターン目のアシスタント応答
1ターン目の指示プロンプトはdeepseek-ai/DeepSeek-V3-0324で合成されています。system promptや2ターン目の指示プロンプトは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-Creative-Writing-39.6k.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Creative-Writing-Sonnet4.6-800x
Pulitzer Diamond Prose CLAUDE Seeds
This dataset contains 833 high-quality creative writing seeds generated using Claude 4.6 Sonnet.
Each entry represents a story opening designed to meet high literary standards.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements: extreme show-don't-tell, double-labor sentence… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-800x.Creative-Writing-Sonnet4.6-Cleaned
Creative-Writing-Sonnet4.6-Cleaned
Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved.
Format
Each line is a JSON object with:
messages: list of message dicts with roles (system, user, assistant)
System: writing quality instructions
User: cleaned creative writing prompt
Assistant: creative writing response (may include <think> traces)
Stats
Metric
Value
Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.Creative-Writing-Part-Two
Creative Writing - Part Two (The Nuclear Dataset)
This dataset represents the "Nuclear" layer of our creative writing training pipeline. While Part One focused on physical and psychological grounding (Shadow & Skeleton), Part Two focuses on dense literary resonance, subtext, and stylistic sophistication.
Methodology: The Nuclear Pipeline
This dataset was built using a multi-phase "Controlled Criticality" approach to ensure maximum signal density without the… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Part-Two.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.CommonCrawl-CreativeCommons-strict
Common Crawl Creative Commons Corpus Strict (C5s)
A filtered version of the Common Crawl Creative Commons Corpus (C5), only retaining samples that:
are also present in the FineWeb or FineWeb-2 datasets;
have no license disagreement (all found licenses have the same type; version number might differ);
are not "non-commercial" ("nc" in license);
are not "cc-unknown";
do not have "wiki" in their name (the idea is that you should include Wikipedia and other Wikidata from other… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/CommonCrawl-CreativeCommons-strict.Creative-Writing-KimiK2.5-Cleaned
Creative-Writing-KimiK2.5-Cleaned
Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved.
Format
Each line is a JSON object with:
messages: list of message dicts with roles (system, user, assistant)
System: writing quality instructions
User: cleaned creative writing prompt
Assistant: creative writing response (may include <think> traces)
Stats
Metric
Value
Total prompt tokens
80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.Creative-Writing-Qwen3.5Plus-2000x
Pulitzer Diamond Prose QWEN Seeds
This dataset contains 2638 high-quality creative writing seeds generated using Qwen 2.5 72B.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Qwen3.5Plus-2000x.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.cqa-creative-writing-expert-cot-preview
CQA: Creative Quality Alignment — Research-Grade Schema v2
English
This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.creative_writing
Dataset Card for telecomadm1145/creative_writing
Dataset Details
Dataset Description
This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output).
The dataset emphasizes:
Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.creative-rubrics-preferences
creative-rubrics-preferences 🎏
A dataset of creative responses using GPT-4.5, o3-mini and DeepSeek-R1.
This dataset contains several prompts seeking creative and diverse answers (like writing movie reviews, short stories, etc), and the style of the responses has been enhanced by prompting the model with custom rubrics that seek different creative styles.
This dataset was used in the paper Configurable Preference Tuning with Rubric-Guided Synthetic Data.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/creative-rubrics-preferences.Creative-Writing-Reasoning-KimiK2.5-600x
Pulitzer Diamond Prose KIMI Seeds
This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1.
Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation.
How it was made
The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.creativemath_fulldeep-creative-writing-zh
Deep Creative Writing Dialogue Dataset (Chinese)
深度文学创作对话数据集
Dataset Description
High-quality Chinese creative writing dialogues covering novel structure, character development, narrative techniques, symbolism, and literary theory.
高质量中文文学创作对话,涵盖小说结构设计、角色塑造、叙事技巧、象征主义、文学理论等议题。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional context (if any)
output: AI response
metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-creative-writing-zh.incremental-instruction-creative-writing
Incremental Instruction Creative Writing
Does delivering a writing brief over several conversation turns change what a
language model writes? This dataset supports that question with matched
creative-writing tasks evaluated under two delivery conditions:
FULL: the complete brief is supplied in one turn.
SHARDED: the same intended brief is introduced across five to nine turns.
The benchmark holds task content fixed while varying how the instructions are
delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.smoltalk-creative-writing-enPurified-openai-messages
📖 SmolTalk-Creative-Writing-enPurified-openai-messages
SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset.
The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.Japanese-Creative-Writing-GLM4.5
Japanese-Creative-Writing-GLM4.5
概要
日本語の小説執筆タスクのデータセットであるAratako/Japanese-Creative-Writing-39.6kから一部の指示を抽出し、zai-org/GLM-4.5で応答を再生成した約8000件のデータセットです。
データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction: 指示プロンプト
output: アシスタント応答
system promptは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス
MITライセンスの元配布します。
figaro-creative-writing
figaro-creative-writing
A high-quality creative writing dataset built using an editor feedback pipeline. Each story goes through three stages: DeepSeek V3.2 writes a first draft, Grok 4.1 Fast provides detailed editorial feedback, then DeepSeek revises based on that feedback. The revision is the final output.
Overview
Rows
6,022
Format
Single-turn chat (system + user + assistant)
Prompts
Gryphe/Opus-WritingPrompts
Writer model
DeepSeek V3.2 (temp=1.0… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/figaro-creative-writing.Opus-4.6-RU-Reasoning-creative-1385x-not-filtered
Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset
A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response.
Dataset Info
Language: Russian 🇷🇺
Size: ~1,385 samples (growing)
Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"}
Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.DesignBench
Dataset Card for DesignBench
Dataset Summary
DesignBench is a multi-framework, multi-task benchmark for evaluating MLLM-based front-end engineering. The paper targets limitations of prior UI code generation benchmarks by covering React, Vue, Angular, and vanilla HTML/CSS, and by evaluating generation, edit, and repair workflows. The full benchmark contains 900 webpage samples spanning multiple topics, edit types, and issue categories.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/creative-graphic-design/DesignBench.creative-rubrics
creative-rubrics 🎏
A dataset of creative responses using GPT-4.5, o3-mini and DeepSeek-R1.
This dataset contains several prompts seeking creative and diverse answers (like writing movie reviews, short stories, etc), and the style of the responses has been enhanced by prompting the model with custom rubrics that seek different creative styles.
It can be used for finetuning for custom styles with open-text tasks.
The dataset was presented in the paper Configurable Preference Tuning… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/creative-rubrics.ptbr-creative-cpt-qwen35-08b-v02
PT-BR Creative CPT — Qwen3.5-0.8B data-prep v0.2
This repository is a derived, model/tokenizer-specific training artifact for continued pretraining experiments.
It is not the canonical text corpus.
Canonical source:
oliveirabruno01/ptbr-creative-cpt
Canonical corpus fingerprint:
21f72f64b3b73425bc78d91046a52aefddb8413b747d69f3422c31da8f536840
Identity
Model/tokenizer: Qwen/Qwen3.5-0.8B-Base
Context length: 2048
Data-prep version: v0.2
Primary split policy:… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/ptbr-creative-cpt-qwen35-08b-v02.Qwill-RP-CreativeWriting-Reasoning
Qwill RP CreativeWriting Reasoning Dataset
📝 Dataset Summary
Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes:
Reasoning, wrapped in <think>...</think>
Final Answer, wrapped in <answer>...</answer>
The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.
