CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stindardlogic /creative-writing-sft-50k Creative Writing SFT (50K) 50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure. Motivation Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.texttext-generation10K<n<100K0 likes5k downloads2mo agoHugging Face02ChaoticNeutrals /Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed. (Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined. Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.text1K<n<10K18 likes2.9k downloads2y agoHugging Face03Dampfinchen /Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course). This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.text1K<n<10K36 likes2.1k downloads8mo agoHugging Face04Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.1k downloads2mo agoHugging Face05Crownelius /Creative-Writing-Gemini3Pro-2700x Pulitzer Diamond Prose GEMINI Seeds This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.texttext-generation1K<n<10K5 likes1.6k downloads2mo agoHugging Face06Crownelius /Creative-Writing-Sonnet4.6-800x Pulitzer Diamond Prose CLAUDE Seeds This dataset contains 833 high-quality creative writing seeds generated using Claude 4.6 Sonnet. Each entry represents a story opening designed to meet high literary standards. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements: extreme show-don't-tell, double-labor sentence… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-800x.texttext-generationn<1K7 likes894 downloads2mo agoHugging Face07Crownelius /Creative-Writing-Sonnet4.6-Cleaned Creative-Writing-Sonnet4.6-Cleaned Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.texttext-generationn<1K3 likes853 downloads2mo agoHugging Face08Samsoup /CreativeEval CreativeEval Paper-grouped multidimensional research-ideation evaluation data from CreativeEval. Contents The release contains 1,026 complete paper rows. Each row contains a human-written research-paper introduction, the raw reviewer score arrays for provenance, and four mean prediction targets: contribution_mean, soundness_mean, presentation_mean, and overall_score_mean. All four targets are derived from the human reviewer scores released with the paper. The raw… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/CreativeEval.text1K<n<10K0 likes763 downloads2mo agoHugging Face09Crownelius /Creative-Writing-Part-Two Creative Writing - Part Two (The Nuclear Dataset) This dataset represents the "Nuclear" layer of our creative writing training pipeline. While Part One focused on physical and psychological grounding (Shadow & Skeleton), Part Two focuses on dense literary resonance, subtext, and stylistic sophistication. Methodology: The Nuclear Pipeline This dataset was built using a multi-phase "Controlled Criticality" approach to ensure maximum signal density without the… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Part-Two.texttext-generation1K<n<10K3 likes600 downloads2mo agoHugging Face10Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes338 downloads2mo agoHugging Face11Crownelius /Creative-Writing-Qwen3.5Plus-2000x Pulitzer Diamond Prose QWEN Seeds This dataset contains 2638 high-quality creative writing seeds generated using Qwen 2.5 72B. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Qwen3.5Plus-2000x.texttext-generation1K<n<10K2 likes328 downloads2mo agoHugging Face12Ttimofeyka /Creative-Writing-Multiturn-Cleaned-16ktext1K<n<10K0 likes276 downloads1y agoHugging Face13BreadStudio /cqa-creative-writing-expert-cot-preview CQA: Creative Quality Alignment — Research-Grade Schema v2 English This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.texttext-generationn<1K6 likes231 downloads2mo agoHugging Face14empathielabs /creative_writing_conversationtext1K<n<10K0 likes223 downloads2y agoHugging Face15telecomadm1145 /creative_writing Dataset Card for telecomadm1145/creative_writing Dataset Details Dataset Description This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output). The dataset emphasizes: Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.texttext-generation1K<n<10K6 likes203 downloads1y agoHugging Face16Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes147 downloads2mo agoHugging Face17DataPilot /Creative-Writing-Dataset Creative Writing Dataset(クリエイティブライティングデータセット) 概要 本データセットは、Aratako/Japanese-Creative-Writing-39.6k の instruction_1 / instruction_2 をそのまま保持し、output_1 / output_2 を Kimi K2.5(Reasoning effort=high) で再生成した ロールプレイング創作データセット です。コンテンツレーティングは R15以下 に制約されています。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約5,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1〜2ターン(instruction + 再生成output) コンテンツレーティング R15以下 ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Creative-Writing-Dataset.text1K<n<10K0 likes140 downloads6mo agoHugging Face18evoeval /EvoEval_creativetextn<1K0 likes134 downloads2y agoHugging Face19Disya /eq-bench-creative-writing-v3https://eqbench.com/creative_writing.html eq-bench textn<1K2 likes114 downloads1y agoHugging Face20AngelWarmSmile123 /deep-creative-writing-zh Deep Creative Writing Dialogue Dataset (Chinese) 深度文学创作对话数据集 Dataset Description High-quality Chinese creative writing dialogues covering novel structure, character development, narrative techniques, symbolism, and literary theory. 高质量中文文学创作对话,涵盖小说结构设计、角色塑造、叙事技巧、象征主义、文学理论等议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-creative-writing-zh.texttext-generation1K<n<10K1 likes114 downloads3mo agoHugging Face21enPurified /smoltalk-creative-writing-enPurified-openai-messages 📖 SmolTalk-Creative-Writing-enPurified-openai-messages SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset. The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.texttext-generation10K<n<100K2 likes107 downloads9mo agoHugging Face22DarkyMan /Opus-4.6-RU-Reasoning-creative-1385x-not-filtered Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response. Dataset Info Language: Russian 🇷🇺 Size: ~1,385 samples (growing) Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"} Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.texttext-generation1K<n<10K3 likes88 downloads6mo agoHugging Face23Crownelius /Qwen3.5-Creative-Reasoning Qwen3.5-Creative-Reasoning Stats Metric Value Total prompt tokens 0 Total completion tokens 318,934 Total tokens 318,934 Total cost $0.50 (USD) Average turns 1.00 Average tool calls 0.00 Average tokens per row 2,327.99 Cost estimated using Qwen3.5 pricing on OpenRouter ($0.26/M input, $1.56/M output) textn<1K5 likes80 downloads2mo agoHugging Face24TeichAI /mistral-small-creative-500x Mistral Small Creative - 500x This is a non-reasoning dataset created using Mistral Small Creative. The dataset is meant for creating distilled versions of Mistral Small Creative by fine-tuning already existing open-source LLMs. This dataset only covers generating short fictional stories. Meant for testing purposes -> to evaluate if creating a larger dataset would be worth it Stats Costs: $ 0.20 (USD) Total tokens (input + output): 704K textn<1K8 likes74 downloads8mo agoHugging Face25open-llm-leaderboard /Aashraf995__Creative-7B-nerd-detailsgated Dataset Card for Evaluation run of Aashraf995/Creative-7B-nerd Dataset automatically created during the evaluation run of model Aashraf995/Creative-7B-nerd The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Aashraf995__Creative-7B-nerd-details.tabular10K<n<100K0 likes55 downloads2y agoHugging Face26marcuscedricridia /Qwill-RP-CreativeWriting-Reasoning Qwill RP CreativeWriting Reasoning Dataset 📝 Dataset Summary Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes: Reasoning, wrapped in <think>...</think> Final Answer, wrapped in <answer>...</answer> The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.tabulartext-generation1K<n<10K8 likes54 downloads1y agoHugging Face27Zethive /CreativeBench CreativeBench HuggingFace Datasets This directory contains two JSONL datasets used in CreativeBench for evaluating the creative problem‑solving capabilities of code models: combinatorial_creativity_dataset_1308.jsonl exploratory_creativity_dataset_551.jsonl Each line in these files is a standalone JSON object (JSONL format). Below we describe the purpose, scale, and field definitions of each dataset to facilitate reuse and re‑hosting on platforms such as HuggingFace Datasets.… See the full description on the dataset page: https://huggingface.co/datasets/Zethive/CreativeBench.text1K<n<10K2 likes54 downloads6mo agoHugging Face28LucidityAI /PIPKIN-Creative-174k PIPKIN 174K Creative The PIPKIN 174K Creative dataset is the largest dataset in the series of evolving, in-the-wild creative datasets. This dataset is inspired by Pygmalion's PIPPA dataset from 2023. Data is collected by exchanging anonymous data for OSS model usage (GLM-5, GLM 4.7, DeepSeek V3.1, Qwen 3.5 397B, Kimi K2.5, etc). This data shows 1.6 billion tokens of chat data. The average amount of input tokens is ~1.2k. The average amount of completion tokens is ~1235. There are… See the full description on the dataset page: https://huggingface.co/datasets/LucidityAI/PIPKIN-Creative-174k.text100K<n<1M1 likes51 downloads6mo agoHugging Face29open-llm-leaderboard /bunnycore__Qandora-2.5-7B-Creative-detailsgated Dataset Card for Evaluation run of bunnycore/Qandora-2.5-7B-Creative Dataset automatically created during the evaluation run of model bunnycore/Qandora-2.5-7B-Creative The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qandora-2.5-7B-Creative-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face30CreativeBuilds /tool-callingtextn<1K1 likes48 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.