CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes6.3k downloads3y agoHugging Face02stindardlogic /creative-writing-sft-50k Creative Writing SFT (50K) 50,000 ShareGPT-format creative writing conversations across 12 literary forms and 25 themes. Written to demonstrate craft — not just competent completion, but genuine literary quality: specific detail, earned emotion, controlled voice, purposeful structure. Motivation Most LLM creative writing training data optimizes for fluency and completion rather than craft. Models learn to produce writing that reads smoothly but relies on clichés… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/creative-writing-sft-50k.texttext-generation10K<n<100K0 likes4.9k downloads2mo agoHugging Face03ChaoticNeutrals /Creative_Writing-ShareGPTOriginal Dataset Sources: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts, https://huggingface.co/datasets/anthracite-org/nopm_claude_writing_fixed. (Thank the original dataset creators for their work.) (Nopm) Claude / (Grphye) ChatGPT-4o Syntheticly generated creative writing set's combined. Update: Used most up to date version of gryphes, chatGPT-4o set, Rejections/Slop Filtered, Min-hash Deduplication using -… See the full description on the dataset page: https://huggingface.co/datasets/ChaoticNeutrals/Creative_Writing-ShareGPT.text1K<n<10K18 likes2.9k downloads2y agoHugging Face04Crownelius /Creative-Writing-High-Quality-1300x Creative Writing - Part One (Shadow & Skeleton) This dataset is designed to train Large Language Models (LLMs) in grounded creative writing by enforcing a "Think-Before-You-Write" methodology. Methodology: Shadow & Skeleton Most creative writing datasets train models to produce "vibes" or "cinematic descriptions" that often lack physical coherence or psychological depth. This dataset takes a different approach: Shadow Prompts: We generated 1,000+ isomorphic… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-High-Quality-1300x.texttext-generation1K<n<10K7 likes2.2k downloads2mo agoHugging Face05Dampfinchen /Creative_Writing_MultiturnUPDATE 2026: Stronger filtering using a very sophisticated filtering script and new data including a very small subset of https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-SFT-01 reasoning for thinking with a custom system prompt attached. This is suitable for both instruct non-thinking and thinking models, as I have added a system prompt for these few samples that use the tags <!think!> and </!think!> (without exclamation marks of course). This is a dataset merge of many, many high… See the full description on the dataset page: https://huggingface.co/datasets/Dampfinchen/Creative_Writing_Multiturn.text1K<n<10K36 likes2.1k downloads8mo agoHugging Face06Crownelius /Creative-Writing-Gemini3Pro-2700x Pulitzer Diamond Prose GEMINI Seeds This dataset contains 2745 high-quality creative writing seeds generated using Gemini 1.5 Pro. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Gemini3Pro-2700x.texttext-generation1K<n<10K5 likes1.7k downloads2mo agoHugging Face07amydeng2000 /CREAKHome page & Original source: https://github.com/yasumasaonoe/creak text10K<n<100K0 likes1.6k downloads4y agoHugging Face08Crownelius /Creative-Writing-Sonnet4.6-800x Pulitzer Diamond Prose CLAUDE Seeds This dataset contains 833 high-quality creative writing seeds generated using Claude 4.6 Sonnet. Each entry represents a story opening designed to meet high literary standards. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements: extreme show-don't-tell, double-labor sentence… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-800x.texttext-generationn<1K7 likes906 downloads2mo agoHugging Face09Crownelius /Creative-Writing-Sonnet4.6-Cleaned Creative-Writing-Sonnet4.6-Cleaned Cleaned creative writing SFT dataset from Sonnet 4.6 (833 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Sonnet4.6-Cleaned.texttext-generationn<1K3 likes884 downloads2mo agoHugging Face10Samsoup /CreativeEval CreativeEval Paper-grouped multidimensional research-ideation evaluation data from CreativeEval. Contents The release contains 1,026 complete paper rows. Each row contains a human-written research-paper introduction, the raw reviewer score arrays for provenance, and four mean prediction targets: contribution_mean, soundness_mean, presentation_mean, and overall_score_mean. All four targets are derived from the human reviewer scores released with the paper. The raw… See the full description on the dataset page: https://huggingface.co/datasets/Samsoup/CreativeEval.text1K<n<10K0 likes762 downloads2mo agoHugging Face11Crownelius /Creative-Writing-Part-Two Creative Writing - Part Two (The Nuclear Dataset) This dataset represents the "Nuclear" layer of our creative writing training pipeline. While Part One focused on physical and psychological grounding (Shadow & Skeleton), Part Two focuses on dense literary resonance, subtext, and stylistic sophistication. Methodology: The Nuclear Pipeline This dataset was built using a multi-phase "Controlled Criticality" approach to ensure maximum signal density without the… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Part-Two.texttext-generation1K<n<10K3 likes631 downloads2mo agoHugging Face12lambdaWalker /creditCardDetectionDS Synthetic Credit Card Dataset Overview This repository contains a synthetic dataset of credit card images designed for training and validating machine learning models, specifically using the YOLOv8 object detection model. The dataset consists of 3000 items for training and 1000 items for validation and testing. Each image is 800x800 in JPG format and is accompanied by the necessary annotation files for YOLOv8 and Hugging Face. Dataset Features Total Items:… See the full description on the dataset page: https://huggingface.co/datasets/lambdaWalker/creditCardDetectionDS.imageobject-detection1K<n<10K1 likes461 downloads2y agoHugging Face13Danny-1223 /CREBench CREBench CREBench is a benchmark for evaluating large language models (LLMs) on cryptographic binary reverse engineering. Paper: arXiv:2604.03750 Code: wangyu-ovo/CREBench Project Page: CREBench Homepage Dataset Description CREBench measures reverse-engineering performance on cryptographic binaries across four evaluation levels: Level Task L1 Algorithm identification L2 Key (and IV) extraction L3 Wrapper-level code reimplementation L4 Flag… See the full description on the dataset page: https://huggingface.co/datasets/Danny-1223/CREBench.textothern<1K2 likes345 downloads2mo agoHugging Face14Crownelius /Creative-Writing-KimiK2.5-Cleaned Creative-Writing-KimiK2.5-Cleaned Cleaned creative writing SFT dataset from Kimi K2.5 (655 samples). Prompts cleaned, thinking traces preserved. Format Each line is a JSON object with: messages: list of message dicts with roles (system, user, assistant) System: writing quality instructions User: cleaned creative writing prompt Assistant: creative writing response (may include <think> traces) Stats Metric Value Total prompt tokens 80… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-KimiK2.5-Cleaned.texttext-generationn<1K8 likes344 downloads2mo agoHugging Face15Crownelius /Creative-Writing-Qwen3.5Plus-2000x Pulitzer Diamond Prose QWEN Seeds This dataset contains 2638 high-quality creative writing seeds generated using Qwen 2.5 72B. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Qwen3.5Plus-2000x.texttext-generation1K<n<10K2 likes333 downloads2mo agoHugging Face16Ttimofeyka /Creative-Writing-Multiturn-Cleaned-16ktext1K<n<10K0 likes296 downloads1y agoHugging Face17BreadStudio /cqa-creative-writing-expert-cot-preview CQA: Creative Quality Alignment — Research-Grade Schema v2 English This is a public preview of Bread Studio's post-training data derived from expert judgments about creative writing. The data is structured for inspection and reuse. The full 104-item Chinese creative-writing expert knowledge-elicitation collection is not released with this repository. This public preview contains the same 4 curated samples as v1, now represented with a more precise and traceable v2… See the full description on the dataset page: https://huggingface.co/datasets/BreadStudio/cqa-creative-writing-expert-cot-preview.texttext-generationn<1K6 likes242 downloads2mo agoHugging Face18empathielabs /creative_writing_conversationtext1K<n<10K0 likes240 downloads2y agoHugging Face19JulianHJR /crest CREST: Cognitive REasoning Steering at Test‑time TL; DR CREST is a training-free test-time steering framework that discovers cognitive heads via simple offline calibration and then rotates activations during decoding to guide the model’s reasoning—preserving norms to avoid per-model hyperparameter tuning. This improves accuracy and reduces tokens across models and datasets. What is CREST? CREST (Cognitive REasoning Steering at Test-time) identifies… See the full description on the dataset page: https://huggingface.co/datasets/JulianHJR/crest.text1K<n<10K0 likes223 downloads3mo agoHugging Face20telecomadm1145 /creative_writing Dataset Card for telecomadm1145/creative_writing Dataset Details Dataset Description This dataset is a small-scale instruction–response dataset focused on creative writing tasks.Each example consists of a prompt (instruction specifying writing style, perspective, tone, etc.) and a response (a story segment or novel-like output). The dataset emphasizes: Creative Writing (light novel style, emotional narrative, dialogue-driven, descriptive prose).… See the full description on the dataset page: https://huggingface.co/datasets/telecomadm1145/creative_writing.texttext-generation1K<n<10K6 likes208 downloads1y agoHugging Face21Crownelius /Creative-Writing-Reasoning-KimiK2.5-600x Pulitzer Diamond Prose KIMI Seeds This dataset contains 655 high-quality creative writing seeds generated using Kimi-v1. Each entry represents a story opening designed to meet high literary standards, including internal thinking traces used during generation. How it was made The data was generated using a custom multi-platform generation engine. Models were prompted with a specialized "Diamond Quality" seed template that enforces strict literary requirements:… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Creative-Writing-Reasoning-KimiK2.5-600x.texttext-generationn<1K8 likes146 downloads2mo agoHugging Face22Disya /eq-bench-creative-writing-v3https://eqbench.com/creative_writing.html eq-bench textn<1K2 likes145 downloads1y agoHugging Face23DataPilot /Creative-Writing-Dataset Creative Writing Dataset(クリエイティブライティングデータセット) 概要 本データセットは、Aratako/Japanese-Creative-Writing-39.6k の instruction_1 / instruction_2 をそのまま保持し、output_1 / output_2 を Kimi K2.5(Reasoning effort=high) で再生成した ロールプレイング創作データセット です。コンテンツレーティングは R15以下 に制約されています。生成にはSDG-LOOMという合成データ生成パイプラインを用いました。(sdg-loom) データの説明 項目 内容 件数 約5,000件 形式 JSONL(1行1JSON) 言語 日本語 ターン数 1〜2ターン(instruction + 再生成output) コンテンツレーティング R15以下 ソースデータセット… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Creative-Writing-Dataset.text1K<n<10K0 likes143 downloads6mo agoHugging Face24evoeval /EvoEval_creativetextn<1K0 likes136 downloads2y agoHugging Face25AngelWarmSmile123 /deep-creative-writing-zh Deep Creative Writing Dialogue Dataset (Chinese) 深度文学创作对话数据集 Dataset Description High-quality Chinese creative writing dialogues covering novel structure, character development, narrative techniques, symbolism, and literary theory. 高质量中文文学创作对话,涵盖小说结构设计、角色塑造、叙事技巧、象征主义、文学理论等议题。 Dataset Structure Format: JSONL (JSON Lines) Fields: instruction: User message / question input: Additional context (if any) output: AI response metadata:… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-creative-writing-zh.texttext-generation1K<n<10K1 likes129 downloads3mo agoHugging Face26DarkyMan /Opus-4.6-RU-Reasoning-creative-1385x-not-filtered Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response. Dataset Info Language: Russian 🇷🇺 Size: ~1,385 samples (growing) Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"} Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.texttext-generation1K<n<10K3 likes107 downloads6mo agoHugging Face27enPurified /smoltalk-creative-writing-enPurified-openai-messages 📖 SmolTalk-Creative-Writing-enPurified-openai-messages SmolTalk-Creative-Writing-enPurified is a highly curated, "prose-first" subset of the original collinear-ai/smoltalk-creative-writing dataset. The enPurified collection is built on a specific philosophy: Specialization. While the ecosystem has plenty of datasets for coding (StackOverflow, StarCoder) and mathematics (GSM8K), high-quality, fluent English prose often gets diluted when mixed with syntax-heavy code or rigid math… See the full description on the dataset page: https://huggingface.co/datasets/enPurified/smoltalk-creative-writing-enPurified-openai-messages.texttext-generation10K<n<100K2 likes104 downloads9mo agoHugging Face28Shikha180224 /dexfluence-indian-creator-index Dexfluence Indian Creator Index Verified Indian influencer dataset across Instagram, YouTube, and TikTok with engagement rates, follower tier, niche classification, and authenticity scores. Dataset summary 141,000+ verified Indian creators indexed across Instagram, YouTube, and TikTok Top 5,000 by follower count included in this Hugging Face mirror (CC-BY 4.0) Each record includes: handle, name, platform, niche, follower count, engagement rate, country… See the full description on the dataset page: https://huggingface.co/datasets/Shikha180224/dexfluence-indian-creator-index.tabulartabular-classification1K<n<10K0 likes101 downloads4mo agoHugging Face29HmyHxy /finance-Knowledge-Credit-Chinesetextn<1K4 likes95 downloads1y agoHugging Face30connections-dev /create_interpret create_interpret — results and raw data Every artefact behind ManyaWadhwa/create_interpret: raw generations, parsed paths, judge caches, scores, and the two analyses built on them. Code, method and write-ups live in the GitHub repo (Task A, Task B); this repo is the data those documents are computed from. Task A — how CREATE creative utility moves with model and sampling temperature, and how diverse 16 samples of one query actually are. 50 instances of wadhma/CREATE x 16 samples… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/create_interpret.tabularn<1K0 likes93 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.