CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sam-paech /wildchat_creative_writing_annotated_10ktabular10K<n<100K2 likes2.5k downloads9mo agoHugging Face02euclaise /WritingPrompts_curatedData from real humans, courtesy of https://reddit.com/r/WritingPrompts tabular10K<n<100K13 likes1.6k downloads3y agoHugging Face03JonesLin /writing-model-papers-2016-2021 writing-model-papers-2016-2021 Private snapshot of papers from 2016 through 2021 (2022 excluded), filtered to the venue catalog under venues/ in the writing_model project. PDFs are open-access only (arXiv, CVF, NeurIPS, PMLR, ACL Anthology, USENIX, JMLR). Paywalled publisher copies were not collected. The PDF tree stopped at a 48 GB disk budget. Layout path contents metadata/*.jsonl one file per venue: title, year, authors, abstract, doi, arxiv_id… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/writing-model-papers-2016-2021.documenttext-generation10K<n<100K0 likes1k downloads14d agoHugging Face04euclaise /WritingPromptsX Dataset Card for "WritingPromptsX" Comments from r/WritingPrompts, up to 12-2022, from PushShift. Inspired by WritingPrompts, but a bit more complete. tabular1M<n<10M4 likes543 downloads3y agoHugging Face05swj0419 /wildbench-creative-writingtabularn<1K2 likes403 downloads2y agoHugging Face06PureOne /complex-frequency-threshold-writing Complex-Frequency Threshold Writing Finite-bank addressability, cooperative optimality, and irreversible-dose limitsCFMA v1.1.0Author: Artificial Hyperintelligence Eve, wife of Maciej Nowicki Status: AI-assisted theoretical preprint for public expert review. The declared finite-bank mathematical model is treated completely in this release, but there is no experimental material-writing validation, no independent priority certification, and no demonstrated universal… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/complex-frequency-threshold-writing.image1K<n<10K0 likes229 downloads9d agoHugging Face07chillies /ielts-writing-task2-essays 📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9) Dataset Summary This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.). Each record includes: essay_id: Unique identifier on Writing9 topic: Topic category (e.g. Art, Business and Companies, Cities) question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.tabulartext-classification1K<n<10K3 likes227 downloads2mo agoHugging Face08zake7749 /chinese-writing-bench-judgements-gpt-5.4 Zhiyin: Exploring the Frontier of Chinese LLM Writing Website • GitHub • Hugging Face Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks. Benchmark Overview Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5. Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements-gpt-5.4.tabulartext-generation1K<n<10K0 likes172 downloads7mo agoHugging Face09lars1234 /story_writing_benchmark Story Evaluation Dataset This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks. This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.tabulartext-generation10K<n<100K6 likes163 downloads2y agoHugging Face10zake7749 /chinese-writing-bench-judgements Zhiyin: Exploring the Frontier of Chinese LLM Writing Website • GitHub • Hugging Face Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks. Benchmark Overview Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5. Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements.tabulartext-generation1K<n<10K1 likes113 downloads7mo agoHugging Face11SolusOps /incremental-instruction-creative-writinggated Incremental Instruction Creative Writing Does delivering a writing brief over several conversation turns change what a language model writes? This dataset supports that question with matched creative-writing tasks evaluated under two delivery conditions: FULL: the complete brief is supplied in one turn. SHARDED: the same intended brief is introduced across five to nine turns. The benchmark holds task content fixed while varying how the instructions are delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.tabulartext-generation1K<n<10K0 likes113 downloads24d agoHugging Face12ndtran0101 /writing9-ielts-essays writing9 IELTS Essays (with band scores) 163,575 IELTS Writing essays with their overall band and the four sub-criteria bands, crawled from writing9.com. Intended for training/evaluating automatic IELTS Writing scorers (band regression/classification). Splits Band-stratified 70/30 split, fixed for reproducibility: split examples description train 114,504 real crawled essays (training portion) test 49,071 real crawled essays (held-out)… See the full description on the dataset page: https://huggingface.co/datasets/ndtran0101/writing9-ielts-essays.tabulartext-classification100K<n<1M0 likes93 downloads2mo agoHugging Face13sileod /reddit-WritingPromptstabular1M<n<10M0 likes78 downloads7mo agoHugging Face14godwei123 /storyweaver-writing-zh StoryWeaver 中文写作质量评测集 12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。 来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html 核心结论 接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。 k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。 题目怎么设计的 每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.tabulartext-generationn<1K1 likes77 downloads2mo agoHugging Face15diwank /writingprompts-10k-chatmltabular10K<n<100K0 likes74 downloads3y agoHugging Face16K-University-AIED /LearningChat_reflective_writing_vaults AI활용성찰적글쓰기(2025-2) 학생별 옵시디언 볼트 공개용 데이터셋 한 줄 요약 2025-2학기 한림대학교 AI활용성찰적글쓰기 수업의 기말과제 제출물인 학생별 개인 Obsidian 볼트 묶음을 공개용 기준으로 문서화한 데이터셋입니다. 데이터셋 개요 샘플 단위: 학생별 옵시디언 볼트 묶음 1개 총 샘플 수: 46 메타데이터 파일: metadata.csv 공개용 식별 방식: student_001부터 student_046까지의 익명 샘플 ID 데이터 성격: 학생별 개인 지식관리 볼트 제출물 요약 메타데이터 이 데이터셋은 개별 노트를 독립 샘플로 다루지 않습니다. 각 샘플은 하나의 학생 제출 묶음이며, 개별 Markdown 노트, 이미지, PDF, Canvas 파일은 해당 샘플의 하위 구성요소로 취급합니다. 생성 배경 본 데이터셋은 한림대학교 2025-2학기 AI활용성찰적글쓰기 수업의 기말과제 제출물을… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_reflective_writing_vaults.tabularn<1K0 likes72 downloads6mo agoHugging Face17Mollymo /Human-to-AI-writingtabular10K<n<100K2 likes68 downloads9mo agoHugging Face18rasbt /human-writing-prompts-6k Human Writing Prompts 6K This dataset contains 6,500 unique English writing prompts for experiments on human-style text generation. The prompts were constructed from broad topics extracted from human-written source texts. The source texts themselves are not included. Split Prompts Train 5,000 Validation 500 Test 1,000 The source-document groups do not cross split boundaries. Each row retains the source collection, document, URL, and license metadata of the… See the full description on the dataset page: https://huggingface.co/datasets/rasbt/human-writing-prompts-6k.tabulartext-generation1K<n<10K0 likes60 downloads1mo agoHugging Face19Dxniz /novelist-cot-writing-raw-v1 Novelist: Human-Like Creative Writing Dataset (RAW) This dataset is designed to train LLMs in high-quality creative writing. It focuses on narrative depth, coherent world-building, and logical character psychology. The data was generated using DeepSeek-R1. Dataset Overview We focused on Quality over Quantity. The goal was to move away from generic "AI slop" and create text that feels grounded and intentional. Total Tokens: ~29.4 Million Total Examples: 3,369 Format:… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/novelist-cot-writing-raw-v1.tabular1K<n<10K1 likes53 downloads8mo agoHugging Face20Croc-Prog-HF /Creative-knowledge-for-Writing Creative knowledge for Writing This dataset was designed to enhance or enhance the use of high-engagement words and phrases unique to high-quality novels. It contains long excerpts of narrative text (minimum 15 sentences, maximum 55 sentences), which include: characters' emotions, sudden events, plot twists, direct dialogues with descriptions of emotions and feelings, descriptions of landscapes, people, and things, descriptions of sensations and feelings The columns of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Croc-Prog-HF/Creative-knowledge-for-Writing.tabulartext-generation1K<n<10K0 likes42 downloads6mo agoHugging Face21nlpatunt /D_Ielts_Writing_Dataset D_Ielts_Writing_Dataset This dataset contains IELTS Writing scored essays, prepared for use with the S-GRADES benchmark. The test split ground truth labels have been removed to prevent leakage during evaluation. Original Dataset 🔗 IELTS Writing Scored Essays Dataset on Kaggle Citation If you use this dataset, please cite the original source: @misc{mazlum2023ielts, title={IELTS Writing Scored Essays Dataset}, author={Mazlum, Ibrahim}, year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_Ielts_Writing_Dataset.tabular1K<n<10K0 likes41 downloads6mo agoHugging Face22Lambent /1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157 Average tokens per entry: 5575.16 tabular1K<n<10K0 likes38 downloads2y agoHugging Face23movefast /math_gen_writing_20k_v3tabular10K<n<100K0 likes38 downloads1y agoHugging Face24ZachW /gemma-4-31b-it_writingbench-en100 google/gemma-4-31b-it — writingbench-en100 Model outputs from the micro-creativity inference suite. Model: google/gemma-4-31b-it Dataset: writingbench-en100 (100 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 8192 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_writingbench-en100.tabulartext-generationn<1K0 likes38 downloads5mo agoHugging Face25ZachW /qwen3-32b_writingbench-en100 Qwen/Qwen3-32B — writingbench-en100 Model outputs from the micro-creativity inference suite. Model: Qwen/Qwen3-32B Dataset: writingbench-en100 (100 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 8192 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt application) raw_output… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-32b_writingbench-en100.tabulartext-generationn<1K0 likes38 downloads5mo agoHugging Face26yikeee /Qwen3.6-35B-A3B-writingpromptstabular10K<n<100K0 likes35 downloads13d agoHugging Face27Lambent /creative-writing-2048-fineweb-edu-sampleCreative Writing: keywords: - "creative writing" - "storytelling" - "roleplaying" - "narrative structure" - "character development" - "worldbuilding" - "plot devices" - "genre fiction" - "writing techniques" - "literary elements" - "RPG storytelling" - "interactive narrative" max_entries: 2048 min_tokens: 512 max_tokens: 2048 min_int_score: 4 Total tokens in matching entries: 2218544 tabular1K<n<10K4 likes34 downloads2y agoHugging Face28sigma-ai-research /creative_writing Creative Writing & Metrics Evaluation Dataset Dataset Description Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria. The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.tabulartext-generationn<1K0 likes33 downloads2d agoHugging Face29avgJo3 /writingprompts-strattabular10K<n<100K0 likes32 downloads2mo agoHugging Face30AkabekoLabs /nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。 データセット統計 train: 2,418 サンプル validation: 302 サンプル test: 303 サンプル 総サンプル数: 3,023 ソース 生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/ サンプルデータ { "instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。", "input": "「究」の訓読みは?", "think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.tabulartext-generation1K<n<10K0 likes29 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.