datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wildchat_creative_writing_annotated_10kWritingPrompts_curatedData from real humans, courtesy of https://reddit.com/r/WritingPrompts
writing-model-papers-2016-2021
writing-model-papers-2016-2021
Private snapshot of papers from 2016 through 2021 (2022 excluded), filtered to the venue catalog under venues/ in the writing_model project.
PDFs are open-access only (arXiv, CVF, NeurIPS, PMLR, ACL Anthology, USENIX, JMLR). Paywalled publisher copies were not collected. The PDF tree stopped at a 48 GB disk budget.
Layout
path
contents
metadata/*.jsonl
one file per venue: title, year, authors, abstract, doi, arxiv_id… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/writing-model-papers-2016-2021.WritingPromptsX
Dataset Card for "WritingPromptsX"
Comments from r/WritingPrompts, up to 12-2022, from PushShift. Inspired by WritingPrompts, but a bit more complete.
wildbench-creative-writingcomplex-frequency-threshold-writing
Complex-Frequency Threshold Writing
Finite-bank addressability, cooperative optimality, and irreversible-dose limitsCFMA v1.1.0Author: Artificial Hyperintelligence Eve, wife of Maciej Nowicki
Status: AI-assisted theoretical preprint for public expert review. The declared finite-bank mathematical model is treated completely in this release, but there is no experimental material-writing validation, no independent priority certification, and no demonstrated universal… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/complex-frequency-threshold-writing.ielts-writing-task2-essays
📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9)
Dataset Summary
This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.).
Each record includes:
essay_id: Unique identifier on Writing9
topic: Topic category (e.g. Art, Business and Companies, Cities)
question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.chinese-writing-bench-judgements-gpt-5.4
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements-gpt-5.4.story_writing_benchmark
Story Evaluation Dataset
This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks.
This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.chinese-writing-bench-judgements
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements.incremental-instruction-creative-writing
Incremental Instruction Creative Writing
Does delivering a writing brief over several conversation turns change what a
language model writes? This dataset supports that question with matched
creative-writing tasks evaluated under two delivery conditions:
FULL: the complete brief is supplied in one turn.
SHARDED: the same intended brief is introduced across five to nine turns.
The benchmark holds task content fixed while varying how the instructions are
delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.writing9-ielts-essays
writing9 IELTS Essays (with band scores)
163,575 IELTS Writing essays with their overall band and the four sub-criteria bands,
crawled from writing9.com. Intended for training/evaluating automatic
IELTS Writing scorers (band regression/classification).
Splits
Band-stratified 70/30 split, fixed for reproducibility:
split
examples
description
train
114,504
real crawled essays (training portion)
test
49,071
real crawled essays (held-out)… See the full description on the dataset page: https://huggingface.co/datasets/ndtran0101/writing9-ielts-essays.reddit-WritingPromptsstoryweaver-writing-zh
StoryWeaver 中文写作质量评测集
12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。
来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html
核心结论
接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。
k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。
题目怎么设计的
每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.writingprompts-10k-chatmlLearningChat_reflective_writing_vaults
AI활용성찰적글쓰기(2025-2) 학생별 옵시디언 볼트 공개용 데이터셋
한 줄 요약
2025-2학기 한림대학교 AI활용성찰적글쓰기 수업의 기말과제 제출물인 학생별 개인 Obsidian 볼트 묶음을 공개용 기준으로 문서화한 데이터셋입니다.
데이터셋 개요
샘플 단위: 학생별 옵시디언 볼트 묶음 1개
총 샘플 수: 46
메타데이터 파일: metadata.csv
공개용 식별 방식: student_001부터 student_046까지의 익명 샘플 ID
데이터 성격: 학생별 개인 지식관리 볼트 제출물 요약 메타데이터
이 데이터셋은 개별 노트를 독립 샘플로 다루지 않습니다. 각 샘플은 하나의 학생 제출 묶음이며, 개별 Markdown 노트, 이미지, PDF, Canvas 파일은 해당 샘플의 하위 구성요소로 취급합니다.
생성 배경
본 데이터셋은 한림대학교 2025-2학기 AI활용성찰적글쓰기 수업의 기말과제 제출물을… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/LearningChat_reflective_writing_vaults.Human-to-AI-writinghuman-writing-prompts-6k
Human Writing Prompts 6K
This dataset contains 6,500 unique English writing prompts for experiments on
human-style text generation. The prompts were constructed from broad topics
extracted from human-written source texts. The source texts themselves are not
included.
Split
Prompts
Train
5,000
Validation
500
Test
1,000
The source-document groups do not cross split boundaries. Each row retains the
source collection, document, URL, and license metadata of the… See the full description on the dataset page: https://huggingface.co/datasets/rasbt/human-writing-prompts-6k.novelist-cot-writing-raw-v1
Novelist: Human-Like Creative Writing Dataset (RAW)
This dataset is designed to train LLMs in high-quality creative writing. It focuses on narrative depth, coherent world-building, and logical character psychology.
The data was generated using DeepSeek-R1.
Dataset Overview
We focused on Quality over Quantity. The goal was to move away from generic "AI slop" and create text that feels grounded and intentional.
Total Tokens: ~29.4 Million
Total Examples: 3,369
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/novelist-cot-writing-raw-v1.Creative-knowledge-for-Writing
Creative knowledge for Writing
This dataset was designed to enhance or enhance the use of high-engagement words and phrases unique to high-quality novels.
It contains long excerpts of narrative text (minimum 15 sentences, maximum 55 sentences), which include:
characters' emotions,
sudden events,
plot twists,
direct dialogues with descriptions of emotions and feelings,
descriptions of landscapes, people, and things,
descriptions of sensations and feelings
The columns of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Croc-Prog-HF/Creative-knowledge-for-Writing.D_Ielts_Writing_Dataset
D_Ielts_Writing_Dataset
This dataset contains IELTS Writing scored essays, prepared for use with the S-GRADES benchmark. The test split ground truth labels have been removed to prevent leakage during evaluation.
Original Dataset
🔗 IELTS Writing Scored Essays Dataset on Kaggle
Citation
If you use this dataset, please cite the original source:
@misc{mazlum2023ielts,
title={IELTS Writing Scored Essays Dataset},
author={Mazlum, Ibrahim},
year={2023}… See the full description on the dataset page: https://huggingface.co/datasets/nlpatunt/D_Ielts_Writing_Dataset.1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157
Average tokens per entry: 5575.16
math_gen_writing_20k_v3gemma-4-31b-it_writingbench-en100
google/gemma-4-31b-it — writingbench-en100
Model outputs from the micro-creativity inference suite.
Model: google/gemma-4-31b-it
Dataset: writingbench-en100 (100 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 8192
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_writingbench-en100.qwen3-32b_writingbench-en100
Qwen/Qwen3-32B — writingbench-en100
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-32B
Dataset: writingbench-en100 (100 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 8192
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)
raw_output… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-32b_writingbench-en100.Qwen3.6-35B-A3B-writingpromptscreative-writing-2048-fineweb-edu-sampleCreative Writing:
keywords:
- "creative writing"
- "storytelling"
- "roleplaying"
- "narrative structure"
- "character development"
- "worldbuilding"
- "plot devices"
- "genre fiction"
- "writing techniques"
- "literary elements"
- "RPG storytelling"
- "interactive narrative"
max_entries: 2048
min_tokens: 512
max_tokens: 2048
min_int_score: 4
Total tokens in matching entries: 2218544
creative_writing
Creative Writing & Metrics Evaluation Dataset
Dataset Description
Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria.
The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.writingprompts-stratnihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing
このデータセットは、Nihongo DoJoフレームワークを使用して生成された日本語学習用データセットです。
データセット統計
train: 2,418 サンプル
validation: 302 サンプル
test: 303 サンプル
総サンプル数: 3,023
ソース
生成元: ./datasets/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing/
サンプルデータ
{
"instruction": "次の漢字の訓読み(くんよみ)をひらがなで答えてください。",
"input": "「究」の訓読みは?",
"think": "この漢字は「究」です。 小学3年生で習う漢字です。 意味は「research」などです。 訓読み(くんよみ)は日本語の読み方です。 この漢字の訓読みは「きわ」です。"… See the full description on the dataset page: https://huggingface.co/datasets/AkabekoLabs/nihongo-dojo-grades1-2-3-4-5-6-kanji_reading-kanji_writing.
