datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
incremental-instruction-creative-writing
Incremental Instruction Creative Writing
Does delivering a writing brief over several conversation turns change what a
language model writes? This dataset supports that question with matched
creative-writing tasks evaluated under two delivery conditions:
FULL: the complete brief is supplied in one turn.
SHARDED: the same intended brief is introduced across five to nine turns.
The benchmark holds task content fixed while varying how the instructions are
delivered. It is… See the full description on the dataset page: https://huggingface.co/datasets/SolusOps/incremental-instruction-creative-writing.Qwill-RP-CreativeWriting-Reasoning
Qwill RP CreativeWriting Reasoning Dataset
📝 Dataset Summary
Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes:
Reasoning, wrapped in <think>...</think>
Final Answer, wrapped in <answer>...</answer>
The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.Creative-knowledge-for-Writing
Creative knowledge for Writing
This dataset was designed to enhance or enhance the use of high-engagement words and phrases unique to high-quality novels.
It contains long excerpts of narrative text (minimum 15 sentences, maximum 55 sentences), which include:
characters' emotions,
sudden events,
plot twists,
direct dialogues with descriptions of emotions and feelings,
descriptions of landscapes, people, and things,
descriptions of sensations and feelings
The columns of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/Croc-Prog-HF/Creative-knowledge-for-Writing.creative_writing
Creative Writing & Metrics Evaluation Dataset
Dataset Description
Each row is one human-written continuation of a creative-writing prompt, scored automatically by four LLM judges (gemini-2.0-flash, gemini-3.8-flash, gpt-4o, gpt-5.6-terra) and a set of traditional NLP metrics, and reviewed independently by multiple human raters on the same criteria.
The dataset consists of responses to creative writing prompts. Each prompt specifically contained a direction to… See the full description on the dataset page: https://huggingface.co/datasets/sigma-ai-research/creative_writing.qwen3-8b_arena-hard-creative-writing
Qwen/Qwen3-8B — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-8B
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt application)… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-8b_arena-hard-creative-writing.qwen3-32b_arena-hard-creative-writing
Qwen/Qwen3-32B — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-32B
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-32b_arena-hard-creative-writing.gemma-3-27b-it_arena-hard-creative-writing
google/gemma-3-27b-it — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_arena-hard-creative-writing.llama-3.1-8b-instruct_arena-hard-creative-writing
meta-llama/Llama-3.1-8B-Instruct — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: meta-llama/Llama-3.1-8B-Instruct
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/llama-3.1-8b-instruct_arena-hard-creative-writing.nanbeige4-3b-thinking-2511_arena-hard-creative-writing
Nanbeige/Nanbeige4-3B-Thinking-2511 — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: Nanbeige/Nanbeige4-3B-Thinking-2511
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/nanbeige4-3b-thinking-2511_arena-hard-creative-writing.qwen3-30b-a3b_arena-hard-creative-writing
Qwen/Qwen3-30B-A3B — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: Qwen/Qwen3-30B-A3B
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/qwen3-30b-a3b_arena-hard-creative-writing.olmo-3-7b-think_arena-hard-creative-writing
allenai/OLMo-3-7B-Think — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Think
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-think_arena-hard-creative-writing.olmo-3-7b-instruct_arena-hard-creative-writing
allenai/OLMo-3-7B-Instruct — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: allenai/OLMo-3-7B-Instruct
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/olmo-3-7b-instruct_arena-hard-creative-writing.mistral-small-3.2-24b-instruct-2506_arena-hard-creative-writing
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_arena-hard-creative-writing.gemma-4-31b-it_arena-hard-creative-writing
google/gemma-4-31b-it — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: google/gemma-4-31b-it
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-4-31b-it_arena-hard-creative-writing.gpt-oss-20b_arena-hard-creative-writing
openai/gpt-oss-20b — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: openai/gpt-oss-20b
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gpt-oss-20b_arena-hard-creative-writing.
