CoolFace
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gryphe /Opus-WritingPrompts Opus Writing Prompts This is a dataset containing 3008 short stories, generated by an unrestrained Claude Opus using Reddit's Writing Prompts as a source. Each sample is generally between 4000-6000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Disclaimer: This dataset is extremely varied and includes erotica. You have been warned. Three files are included: A ShareGPT dataset, ready to be used for… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/Opus-WritingPrompts.texttext-generation1K<n<10K86 likes7.6k downloads2y agoHugging Face02Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes7.2k downloads2y agoHugging Face03stindardlogic /creative-writing-sft-50k Creative Writing SFT (50K) 50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure. Motivation Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.texttext-generation10K<n<100K0 likes5.1k downloads2mo agoHugging Face04Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.1k downloads2mo agoHugging Face05Crownelius /Creative-Writing-Gemini3Pro-2700x Pulitzer Diamond Prose GEMINI Seeds This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.texttext-generation1K<n<10K5 likes1.7k downloads2mo agoHugging Face06Crownelius /Creative-Writing-Sonnet4.6-800x Pulitzer Diamond Prose CLAUDE Seeds This dataset contains 833 high-quality creative writing seeds generated using Claude 4.6 Sonnet. Each entry represents a story opening designed to meet high literary standards. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements: extreme show-don't-tell, double-labor sentence… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-800x.texttext-generationn<1K7 likes902 downloads2mo agoHugging Face07Crownelius /Creative-Writing-Sonnet4.6-Cleaned Creative-Writing-Sonnet4.6-Cleaned Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.texttext-generationn<1K3 likes866 downloads2mo agoHugging Face08Crownelius /Creative-Writing-Part-Two Creative Writing - Part Two (The Nuclear Dataset) This dataset represents the "Nuclear" layer of our creative writing training pipeline. While Part One focused on physical and psychological grounding (Shadow & Skeleton), Part Two focuses on dense literary resonance, subtext, and stylistic sophistication. Methodology: The Nuclear Pipeline This dataset was built using a multi-phase "Controlled Criticality" approach to ensure maximum signal density without the… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Part-Two.texttext-generation1K<n<10K3 likes590 downloads2mo agoHugging Face09Nopm /Opus_WritingStruct Opus Writing Instruct 6k Synthetically generated creative writing data using Claude 3 Opus, by Anthropic, filtered and cleaned using automated means. Focus was placed on having as many genres as possible represented in the data, and to have Claude more openly use its excellent prose. It also contains question-answer instruction pairs related to the topic of writing. Dataset Details Curated by: Nopm License: Apache 2 Credits: The entire SillyTilly community for providing… See the full description on the dataset page: https://huggingface.co/datasets/Nopm/Opus_WritingStruct.texttext-generation1K<n<10K39 likes568 downloads2y agoHugging Face10Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes345 downloads2mo agoHugging Face11Crownelius /Creative-Writing-Qwen3.5Plus-2000x Pulitzer Diamond Prose QWEN Seeds This dataset contains 2638 high-quality creative writing seeds generated using Qwen 2.5 72B. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Qwen3.5Plus-2000x.texttext-generation1K<n<10K2 likes327 downloads2mo agoHugging Face12stindardlogic /writing-quality-dpo-100k Writing Quality DPO (100K) 100,000 DPO preference pairs training models to write with clarity, concision, structure, and impact. Each chosen response demonstrates high-quality prose; each rejected response contains exactly one identified writing defect. Motivation Writing assistance is the #1 use case for LLMs, yet most training data optimizes for factual correctness rather than writing craft. This dataset trains models to distinguish genuinely good writing from… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/writing-quality-dpo-100k.texttext-generation100K<n<1M0 likes311 downloads2mo agoHugging Face13stindardlogic /email-writing-sft-100k Email Writing SFT (100K) 100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails. Motivation Email is the primary communication channel for most professional work, yet LLMs often produce emails that are: Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.texttext-generation100K<n<1M3 likes243 downloads2mo agoHugging Face14BreadStudio /cqa-creative-writing-expert-cot-preview CQA: Creative Quality Alignment — Research-Grade Schema v2 English This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.texttext-generationn<1K6 likes225 downloads2mo agoHugging Face15PinkPixel /Story-Writing 📖 Story-Writing Dataset ✨ This dataset is a collection of creative writing stories based on the Writing Prompts ([WP]) format. It is designed to help models learn how to write compelling, structured, and emotionally engaging narratives. 📂 Dataset Structure The data is provided in ChatML format, making it ideal for instruction tuning. Files writing_train_chatml.jsonl: Training data. writing_valid_chatml.jsonl: Validation data. Example Entry {… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Story-Writing.texttext-generation1M<n<10M3 likes221 downloads5mo agoHugging Face16telecomadm1145 /creative_writing Dataset Card for telecomadm1145/creative_writing Dataset Details Dataset Description This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output). The dataset emphasizes: Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.texttext-generation1K<n<10K6 likes198 downloads1y agoHugging Face17stindardlogic /technical-writing-sft-100k Technical Writing SFT (100K) 100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use. Motivation Technical writing is one of the most underserved capabilities in LLMs. Common model failures: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.texttext-generation100K<n<1M0 likes193 downloads2mo agoHugging Face18Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes150 downloads2mo agoHugging Face19AngelWarmSmile123 /deep-creative-writing-zh Deep Creative Writing Dialogue Dataset (Chinese) 深度文学创作对话数据集 Dataset Description High-quality Chinese creative writing dialogues covering novel structure, character development, narrative techniques, symbolism, and literary theory. 高质量中文文学创作对话,涵盖小说结构设计、角色塑造、叙事技巧、象征主义、文学理论等议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-creative-writing-zh.texttext-generation1K<n<10K1 likes116 downloads3mo agoHugging Face20enPurified /smoltalk-creative-writing-enPurified-openai-messages 📖 SmolTalk-Creative-Writing-enPurified-openai-messages SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset. The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.texttext-generation10K<n<100K2 likes115 downloads9mo agoHugging Face21godwei123 /storyweaver-writing-zh StoryWeaver 中文写作质量评测集 12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。 来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html 核心结论 接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。 k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。 题目怎么设计的 每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.tabulartext-generationn<1K1 likes77 downloads2mo agoHugging Face22PinkPixel /Childrens-Story-Writing 🧒 Children's Story Writing Dataset ✨ This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences). 📂 Dataset Structure The data is provided in ChatML format, making it ideal for instruction tuning. Files writing_train_children.jsonl: Training data. writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.texttext-generation1M<n<10M2 likes67 downloads5mo agoHugging Face23zorqelis-ai /soreqen-writing SoreQen Writing Professional English writing: cold and internal emails, blog posts, and rewriting tasks, most with the brief that produced them retained as a system prompt. Curated and published by ZorQelis AI. Train rows 23,728 Validation rows 484 Format chat messages (JSONL) Format {"messages": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}], "source": "orca_math", "lang": "english"… See the full description on the dataset page: https://huggingface.co/datasets/zorqelis-ai/soreqen-writing.texttext-generation10K<n<100K1 likes43 downloads1mo agoHugging Face24sigma-ai-research /creative_writing Creative Writing & Metrics Evaluation Dataset Dataset Description Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria. The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.tabulartext-generationn<1K0 likes33 downloads2d agoHugging Face25PJMixers-Dev /Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT Gryphe/ChatGPT-4o-Writing-Prompts with responses regenerated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generation1K<n<10K0 likes18 downloads2y agoHugging Face26agentlans /Qwen3.5-9B-Writing-DPO-short-stories Qwen3.5-9B-Writing-DPO Short Stories This is a replication of the agentlans/llama3.1-8b-short-stories dataset using the nbeerbower/Qwen3.5-9B-Writing-DPO experimental writing model. It shows the distinctive tone and style of the model: Long outputs - even longer than requested in the prompt Dense sensory details Complex sentences with varying lengths Very little overt action and movement ("show don't tell" taken to an extreme) Unimaginative character names like "Elara", "Silas"… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/Qwen3.5-9B-Writing-DPO-short-stories.tabulartext-generation1K<n<10K0 likes18 downloads6mo agoHugging Face27agentlans /euclaise-WritingPromptsX WritingPromptsX Filtered Dataset This dataset contains a filtered subset of euclaise/WritingPromptsX, originally collected from Reddit’s r/WritingPrompts using PushShift (up to Dec 2022). It includes the first 100 000 entries where the text length is between 1 000 and 8 000 characters. The data has been cleaned and shuffled. It’s well-suited for creative writing, story generation, and language modeling tasks involving longer text sequences. Data Fields Name… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/euclaise-WritingPromptsX.tabulartext-generation100K<n<1M0 likes15 downloads1y agoHugging Face28agentlans /introspective-writing-prompts Introspective Writing Prompts A dataset of writing prompts designed to reveal a writer’s personality traits and natural writing style. Useful as an introspective exercise for both humans and AI. Each chatbot—Grok, Gemini, Claude, KimiK2 (via HuggingChat), Deepseek, MetaAI, Perplexity, LeChat, ChatGPT, and Copilot—generated prompts following the same base template. Exact duplicates were removed. Prompt generation template Create a diverse list of writing prompts designed… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/introspective-writing-prompts.texttext-generation1K<n<10K0 likes15 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.