CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01godwei123 /storyweaver-writing-zh StoryWeaver 中文写作质量评测集 12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。 来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html 核心结论 接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。 k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。 题目怎么设计的 每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.tabulartext-generationn<1K1 likes75 downloads2mo agoHugging Face02ndtran0101 /writing9-ielts-essays writing9 IELTS Essays (with band scores) 163,575 IELTS Writing essays with their overall band and the four sub-criteria bands, crawled from writing9.com. Intended for training/evaluating automatic IELTS Writing scorers (band regression/classification). Splits Band-stratified 70/30 split, fixed for reproducibility: split examples description train 114,504 real crawled essays (training portion) test 49,071 real crawled essays (held-out)… See the full description on the dataset page: https://huggingface.co/datasets/ndtran0101/writing9-ielts-essays.tabulartext-classification100K<n<1M0 likes67 downloads2mo agoHugging Face03Dxniz /novelist-cot-writing-raw-v1 Novelist: Human-Like Creative Writing Dataset (RAW) This dataset is designed to train LLMs in high-quality creative writing. It focuses on narrative depth, coherent world-building, and logical character psychology. The data was generated using DeepSeek-R1. Dataset Overview We focused on Quality over Quantity. The goal was to move away from generic "AI slop" and create text that feels grounded and intentional. Total Tokens: ~29.4 Million Total Examples: 3,369 Format:… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/novelist-cot-writing-raw-v1.tabular1K<n<10K1 likes49 downloads8mo agoHugging Face04Lambent /1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157 Average tokens per entry: 5575.16 tabular1K<n<10K0 likes39 downloads2y agoHugging Face05Lambent /creative-writing-2048-fineweb-edu-sampleCreative Writing: keywords: - "creative writing" - "storytelling" - "roleplaying" - "narrative structure" - "character development" - "worldbuilding" - "plot devices" - "genre fiction" - "writing techniques" - "literary elements" - "RPG storytelling" - "interactive narrative" max_entries: 2048 min_tokens: 512 max_tokens: 2048 min_int_score: 4 Total tokens in matching entries: 2218544 tabular1K<n<10K4 likes36 downloads2y agoHugging Face06yikeee /Qwen3.6-35B-A3B-writingpromptstabular10K<n<100K0 likes36 downloads13d agoHugging Face07sigma-ai-research /creative_writing Creative Writing & Metrics Evaluation Dataset Dataset Description Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria. The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.tabulartext-generationn<1K0 likes28 downloads1d agoHugging Face08agentlans /euclaise-WritingPromptsX WritingPromptsX Filtered Dataset This dataset contains a filtered subset of euclaise/WritingPromptsX, originally collected from Reddit’s r/WritingPrompts using PushShift (up to Dec 2022). It includes the first 100 000 entries where the text length is between 1 000 and 8 000 characters. The data has been cleaned and shuffled. It’s well-suited for creative writing, story generation, and language modeling tasks involving longer text sequences. Data Fields Name… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/euclaise-WritingPromptsX.tabulartext-generation100K<n<1M0 likes20 downloads1y agoHugging Face09agentlans /Qwen3.5-9B-Writing-DPO-short-stories Qwen3.5-9B-Writing-DPO Short Stories This is a replication of the agentlans/llama3.1-8b-short-stories dataset using the nbeerbower/Qwen3.5-9B-Writing-DPO experimental writing model. It shows the distinctive tone and style of the model: Long outputs - even longer than requested in the prompt Dense sensory details Complex sentences with varying lengths Very little overt action and movement ("show don't tell" taken to an extreme) Unimaginative character names like "Elara", "Silas"… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/Qwen3.5-9B-Writing-DPO-short-stories.tabulartext-generation1K<n<10K0 likes18 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.