CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stindardlogic /creative-writing-sft-50k Creative Writing SFT (50K) 50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure. Motivation Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.texttext-generation10K<n<100K0 likes5.1k downloads2mo agoHugging Face02ChaoticNeutrals /Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed. (Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined. Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.text1K<n<10K18 likes3k downloads2y agoHugging Face03Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.1k downloads2mo agoHugging Face04Dampfinchen /Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course). This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.text1K<n<10K36 likes2.1k downloads8mo agoHugging Face05Crownelius /Creative-Writing-Gemini3Pro-2700x Pulitzer Diamond Prose GEMINI Seeds This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.texttext-generation1K<n<10K5 likes1.7k downloads2mo agoHugging Face06Crownelius /Creative-Writing-Sonnet4.6-800x Pulitzer Diamond Prose CLAUDE Seeds This dataset contains 833 high-quality creative writing seeds generated using Claude 4.6 Sonnet. Each entry represents a story opening designed to meet high literary standards. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements: extreme show-don't-tell, double-labor sentence… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-800x.texttext-generationn<1K7 likes902 downloads2mo agoHugging Face07Crownelius /Creative-Writing-Sonnet4.6-Cleaned Creative-Writing-Sonnet4.6-Cleaned Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.texttext-generationn<1K3 likes866 downloads2mo agoHugging Face08Crownelius /Creative-Writing-Part-Two Creative Writing - Part Two (The Nuclear Dataset) This dataset represents the "Nuclear" layer of our creative writing training pipeline. While Part One focused on physical and psychological grounding (Shadow & Skeleton), Part Two focuses on dense literary resonance, subtext, and stylistic sophistication. Methodology: The Nuclear Pipeline This dataset was built using a multi-phase "Controlled Criticality" approach to ensure maximum signal density without the… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Part-Two.texttext-generation1K<n<10K3 likes590 downloads2mo agoHugging Face09Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes345 downloads2mo agoHugging Face10Crownelius /Creative-Writing-Qwen3.5Plus-2000x Pulitzer Diamond Prose QWEN Seeds This dataset contains 2638 high-quality creative writing seeds generated using Qwen 2.5 72B. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Qwen3.5Plus-2000x.texttext-generation1K<n<10K2 likes327 downloads2mo agoHugging Face11Ttimofeyka /Creative-Writing-Multiturn-Cleaned-16ktext1K<n<10K0 likes273 downloads1y agoHugging Face12BreadStudio /cqa-creative-writing-expert-cot-preview CQA: Creative Quality Alignment — Research-Grade Schema v2 English This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.texttext-generationn<1K6 likes225 downloads2mo agoHugging Face13empathielabs /creative_writing_conversationtext1K<n<10K0 likes214 downloads2y agoHugging Face14telecomadm1145 /creative_writing Dataset Card for telecomadm1145/creative_writing Dataset Details Dataset Description This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output). The dataset emphasizes: Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.texttext-generation1K<n<10K6 likes198 downloads1y agoHugging Face15Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes150 downloads2mo agoHugging Face16DataPilot /Creative-Writing-Dataset Creative Writing Dataset(クリエイティブライティングデータセット) 概要 本データセットは、Aratako/Japanese-Creative-Writing-39.6k の instruction_1 / instruction_2 をそのまま保持し、output_1 / output_2 を Kimi K2.5(Reasoning effort=high) で再生成した ロールプレイング創作データセット です。コンテンツレーティングは R15以下 に制約されています。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約5,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1〜2ターン(instruction + 再生成output) コンテンツレーティング R15以下 ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Creative-Writing-Dataset.text1K<n<10K0 likes126 downloads6mo agoHugging Face17AngelWarmSmile123 /deep-creative-writing-zh Deep Creative Writing Dialogue Dataset (Chinese) 深度文学创作对话数据集 Dataset Description High-quality Chinese creative writing dialogues covering novel structure, character development, narrative techniques, symbolism, and literary theory. 高质量中文文学创作对话,涵盖小说结构设计、角色塑造、叙事技巧、象征主义、文学理论等议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-creative-writing-zh.texttext-generation1K<n<10K1 likes116 downloads3mo agoHugging Face18enPurified /smoltalk-creative-writing-enPurified-openai-messages 📖 SmolTalk-Creative-Writing-enPurified-openai-messages SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset. The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.texttext-generation10K<n<100K2 likes115 downloads9mo agoHugging Face19Disya /eq-bench-creative-writing-v3https://eqbench.com/creative_writing.html eq-bench textn<1K2 likes113 downloads1y agoHugging Face20marcuscedricridia /Qwill-RP-CreativeWriting-Reasoning Qwill RP CreativeWriting Reasoning Dataset 📝 Dataset Summary Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes: Reasoning, wrapped in <think>...</think> Final Answer, wrapped in <answer>...</answer> The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.tabulartext-generation1K<n<10K8 likes53 downloads1y agoHugging Face21Lambent /1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157 Average tokens per entry: 5575.16 tabular1K<n<10K0 likes38 downloads2y agoHugging Face22theprint /CreativeWriting-3ktexttext-generation1K<n<10K1 likes38 downloads10mo agoHugging Face23zerofata /Instruct-Anime-CreativeWritingA small synthetic dataset of instruct prompts from various fandoms, answered by anime / game characters instructed to respond in a creative writing style. Data was created by a mix of Claude 3.7, Deepseek-v3 & Gemini Flash. Dataset has been cleaned for major slop. There's a decent amount of one word (testaments, fractures etc.) slop in here, but generally this one is pretty clean and has a positive impact on writing style for models it's applied to. Creation Process: Character is randomly… See the full description on the dataset page: https://huggingface.co/datasets/zerofata/Instruct-Anime-CreativeWriting.text1K<n<10K11 likes37 downloads1y agoHugging Face24Lambent /creative-writing-2048-fineweb-edu-sampleCreative Writing: keywords: - "creative writing" - "storytelling" - "roleplaying" - "narrative structure" - "character development" - "worldbuilding" - "plot devices" - "genre fiction" - "writing techniques" - "literary elements" - "RPG storytelling" - "interactive narrative" max_entries: 2048 min_tokens: 512 max_tokens: 2048 min_int_score: 4 Total tokens in matching entries: 2218544 tabular1K<n<10K4 likes34 downloads2y agoHugging Face25sigma-ai-research /creative_writing Creative Writing & Metrics Evaluation Dataset Dataset Description Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria. The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.tabulartext-generationn<1K0 likes33 downloads2d agoHugging Face26rx1lora /tb00-creative-writing-uncensored Creative Writing Instruction Dataset 400+ creative writing samples across genres: thriller, romance, sci-fi, horror, literary fiction. Uncensored, high quality. Stats Samples: 78 Format: JSONL (messages format, ready for SFT) License: Apache 2.0 Usage from datasets import load_dataset ds = load_dataset("paijo77/creative-writing-uncensored") # Format: messages array print(ds['train'][0]['messages']) Fine-tuning from trl import… See the full description on the dataset page: https://huggingface.co/datasets/rx1lora/tb00-creative-writing-uncensored.textn<1K0 likes26 downloads1mo agoHugging Face27Ttimofeyka /Creative_Writing_Multiturn_llama3.2text1K<n<10K0 likes19 downloads1y agoHugging Face28Gale0302 /Creativewritingtext1K<n<10K0 likes14 downloads1y agoHugging Face29oyi77 /creative-writing-uncensored Creative Writing Instruction Dataset 400+ creative writing samples across genres: thriller, romance, sci-fi, horror, literary fiction. Uncensored, high quality. Stats Samples: 78 Format: JSONL (messages format, ready for SFT) License: Apache 2.0 Usage from datasets import load_dataset ds = load_dataset("paijo77/creative-writing-uncensored") # Format: messages array print(ds['train'][0]['messages']) Fine-tuning from trl import SFTTrainer #… See the full description on the dataset page: https://huggingface.co/datasets/oyi77/creative-writing-uncensored.textn<1K0 likes11 downloads6mo agoHugging Face30marcuscedricridia /gemm3-creativewritingtextn<1K1 likes10 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.