datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
story_clozestory_writing_benchmark
Story Evaluation Dataset
This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks.
This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.Danbooru-Dataset-csv
Danbooru Dataset CSV
面向 Danbooru 标签管理 / 打标工具的公开元数据合集。这里只放整理后的 CSV,不含任何图片。后续还会继续补充 artist、copyright 等更多表;本页只做项目总览,各文件以仓库里的 CSV 为准。
标签与 wiki 来自 Danbooru。本仓库整理表使用 MIT 协议。原图版权仍归各自作者。
当前文件
文件
内容
截止日期
行数
danbooru_dataset_general_260820.csv
general 通用标签(别名、层级、父子、分类、wiki)
2026-08-20
106,414
danbooru_character_tags.csv
character 角色标签(别名、作品、父标签、投稿数)
2026-07-20
329,747
danbooru_artist_tags.csv
artist 画师标签(译名、数据量)
—
576,842
tag-near-synonym-relations4.csv… See the full description on the dataset page: https://huggingface.co/datasets/StoryAura/Danbooru-Dataset-csv.StorySeedStorySeed is a data set specially designed for training and evaluating the performance of text generation models in the domain of children’s picture book creation. It contains 4376 thoughtfully curated prompt-response pairs, encompassing nine major thematic categories: educational, emotional intelligence and social skills, adventure tales, natural science, folk tales and myths, daily life, humorous stories, bedtime stories, as well as other general picture book stories not specific to any… See the full description on the dataset page: https://huggingface.co/datasets/Aiwensile2/StorySeed.story_analogyStoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding
This is the StoryAnalogy dataset in the EMNLP'23 paper: StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding.
If you use this research, please cite us:
@inproceedings{jiayang2023storyanalogy,
title={StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical… See the full description on the dataset page: https://huggingface.co/datasets/JoeyCheng/story_analogy.kikamba-story-katiwa-cc0devto-war-story-performance
Dev.to War-Story Performance Dataset
941 articles published on dev.to under the @whoffagents account, spanning April 2026. Includes title, tags, engagement metrics, and reading time.
Why this exists
We run an agentic content pipeline that publishes developer war-stories daily. This dataset captures real performance data across article formats to answer: what titles and tags actually get reactions on dev.to?
Preliminary finding: war-story framing ("I did X and here's what… See the full description on the dataset page: https://huggingface.co/datasets/WH0FF/devto-war-story-performance.balinese-story-texts-extendsWorkflow
StorySculptor_datasetstory_emotion_classificationchico_prompts_generate_story
Dataset Card for Dataset Name
This dataset provides around 8,000 prompts in Spanish about short stories.
The following is the prompt in english:
prompt:
Write a short story based on the following title:
{{titles}}
completion:
{{contents}}
In spanish:
prompt:
Escribe una historia corta basada en el siguiente título {{titles}}
completion:
{{contents}}
More Information
This dataset is a sub-version of the original chico dataset.
storySynthetic_Story_Generated_By_Geminipart1-space-warfare-short-story-3pov-sentences-datasetHorror_Story_Generationpart1-space-warfare-short-story-3pov-1bc-datasetred_team_agent_analysis_multilingual_story_analysis_detailed
red_team_agent_analysis_multilingual_story_analysis_detailed
This dataset was automatically uploaded from the red-team-agent repository.
Dataset Information
Original file: multilingual_story_analysis_detailed.csv
Source path: /home/ubuntu/red-team-agent/red_team_agent/analysis/multilingual_story_analysis_detailed.csv
Validation: Valid CSV with 1000 rows, 12 columns (0.1MB)
Usage
import pandas as pd
from datasets import load_dataset
# Load using datasets… See the full description on the dataset page: https://huggingface.co/datasets/aq1048576/red_team_agent_analysis_multilingual_story_analysis_detailed.StoryMakerml-emoji-storystory_emotion_inferencestoryseeker
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Citation
If you use our data, codebook, or models, please cite the following preprint:
Where do people tell stories online? Story Detection Across Online CommunitiesMaria Antoniak, Joel Mire, Maarten Sap, Elliott Ash, Andrew Piper
Quick Start with Colab
You can view a demonstration of how to load our annotations, fetch… See the full description on the dataset page: https://huggingface.co/datasets/mariaantoniak/storyseeker.Short-Storypart1-coherence-intro-space-warfare-short-story-3pov-1bc-datasetstory-angle-benchmarks
Forbes Story Angle Generator Benchmarks
Benchmark dataset of 20 brand story angle cases with individual scores for story angle quality, editorial fit, brand authority, media attention, narrative strength, and SEO visibility.
Built by ForbesPlacement.com.
Dataset Description
This dataset contains benchmark data for an AI-powered platform helping businesses, founders, executives, and marketing teams develop compelling editorial story ideas for premium business… See the full description on the dataset page: https://huggingface.co/datasets/forbes-placement/story-angle-benchmarks.reddit_story_gen_1bhl_copyright_statuses
Biodiversity Heritage Library Copyright Statuses
This dataset contains all unique copyright statuses present in the items.txt.gz file of the Biodiversity Heritage Library open dataset on AWS Open Data. The unique copyright statuses were extracted, grouped and sorted by frequency using the following DuckDB query:
COPY (SELECT CopyrightStatus, COUNT(*) as Count FROM read_csv('https://bhl-open-data.s3.amazonaws.com/data/item.txt.gz') GROUP BY CopyrightStatus ORDER BY Count DESC) TO… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/bhl_copyright_statuses.StoryTellerstory_summaryStoryArcsAnalysis
StoryArcsAnalysis
tags: clustering, narrative complexity, viewer retention
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'StoryArcsAnalysis' dataset is designed to assist Machine Learning practitioners in understanding how different storytelling techniques and narrative complexities influence viewer retention across TV shows. The dataset includes excerpts from TV show scripts, character development arcs, and pivotal plot… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/StoryArcsAnalysis.part1-augment-space-warfare-short-story-3pov-1bc-dataset
