datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
writingprompts
Dataset Card for "writingprompts"
WritingPrompts dataset, as used in Hierarchical Neural Story Generation. Parsed from the archive
RL-Claude-Creative-Writing-SFT
RL-Claude-Creative-Writing-SFT
Alpaca-format dataset. Columns: instruction, input, output
from datasets import load_dataset
ds = load_dataset("SLoonker/RL-Claude-Creative-Writing-SFT", split="train")
wildchat_creative_writing_annotated_10ksmoltalk-creative-writingWritingPrompts_curatedData from real humans, courtesy of https://reddit.com/r/WritingPrompts
WritingPrompts_preferences
Dataset Card for "WritingPrompts_preferences"
Human preference data from r/WritingPrompts
Japanese-Creative-Writing-39.6k
Japanese-Creative-Writing-39.6k
概要
deepseek-ai/DeepSeek-V3-0324を用いて作成した、約39600件の日本語の小説執筆タスクデータセットです。
全てのデータは2ターンのデータとなっています。また、データセット中の一部データはNSFW表現を含みます。
データの詳細
各データは以下のキーを含みます。
messages: OpenAI messages形式の対話データ
instruction_1: 1ターン目の指示プロンプト
output_1: 1ターン目のアシスタント応答
instruction_2: 2ターン目の指示プロンプト
output_2: 2ターン目のアシスタント応答
1ターン目の指示プロンプトはdeepseek-ai/DeepSeek-V3-0324で合成されています。system promptや2ターン目の指示プロンプトは事前に用意した複数種類からランダムに選択されたものが設定されています。
ライセンス… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Japanese-Creative-Writing-39.6k.RL-Claude-Creative-Writing-DPO
RL-Claude-Creative-Writing-DPO
Alpaca-format dataset. Columns: instruction, input, output, rejected
from datasets import load_dataset
ds = load_dataset("SLoonker/RL-Claude-Creative-Writing-DPO", split="train")
dolly_creative_writing
Dataset Card for "dolly_creative_writing"
More Information needed
WritingPromptsX
Dataset Card for "WritingPromptsX"
Comments from r/WritingPrompts, up to 12-2022, from PushShift. Inspired by WritingPrompts, but a bit more complete.
WritingPrompts_binarizedWritingPrompts_preferences, but processed like SHP
wildbench-creative-writingwildchat-writing-1k-sft-bestptb-writingbenchreddit_creepypasta_writingchinese-writing-benchmark
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-benchmark.ielts-writing-task2-essays
📚 IELTS Writing Task 2 Essays & Feedback Dataset (Writing9)
Dataset Summary
This dataset contains 8,000+ real IELTS Writing Task 2 essays crawled from Writing9. It covers 128 real IELTS exam questions categorized into 25 topics (such as Art, Business, Education, Technology, Environment, Government, Health, etc.).
Each record includes:
essay_id: Unique identifier on Writing9
topic: Topic category (e.g. Art, Business and Companies, Cities)
question: Cleaned IELTS… See the full description on the dataset page: https://huggingface.co/datasets/chillies/ielts-writing-task2-essays.Gryphe_ChatGPT_4o_Writing_Prompts_Chinesecreative_writingchinese-writing-bench-judgements
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements.wildchat-creative-writing-3k-rftcreative_writing_v9WritingPrompts-Filtered
WritingPrompts Filtered Dataset (LitBench Decontaminated)
Dataset Description
This dataset contains filtered and decontaminated WritingPrompts from Reddit, specifically processed to remove any overlap with the LitBench test set. This ensures clean training data for language models without test set contamination.
Processing Statistics
Generated: 2025-09-12
Dataset Size
Original dataset: 265174 entries
After decontamination: 199,248 entries… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF/WritingPrompts-Filtered.chinese-writing-bench-judgements-gpt-5.4
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements-gpt-5.4.airoboros_writing_instructions_gpt-4o-miniwildchat-creative-writing-3k-prefOpus-4.5-WritingStyle-1000x-formatted-fixed
Opus-4.5-WritingStyle-1000x-formatted-fixed
Stats
Metric
Value
Total prompt tokens
24,176
Total completion tokens
2,765,527
Total tokens
2,789,703
Total cost
$69.26 (USD)
Average turns
1.00
Average tool calls
0.00
Average tokens per row
549.59
Cost estimated using Claude Opus 4.5 pricing on OpenRouter ($5.0/M input, $25.0/M output)
bilkent-turkish-writings-dataset
Compilation of Bilkent Turkish Writings Dataset
Dataset Description
This is a comprehensive compilation of Turkish creative writings from Bilkent University's Turkish 101 and Turkish 102 courses (2014-2025). The dataset contains 9119 student writings originally created by students and instructors, focusing on creativity, content, composition, grammar, spelling, and punctuation development.
Note: This dataset is a compilation and digitization of publicly available writings… See the full description on the dataset page: https://huggingface.co/datasets/selimfirat/bilkent-turkish-writings-dataset.writing-prompts
Dataset Card for "WritingPrompts"
More Information needed
essayforum_raw_writing_10k
