datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
datacomp200m
Datacomp200m
This is a smaller version of the datacomp_1b dataset.
Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows.
The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling.
Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.LoC-PD-Books
Library of Congress Public Domain Books (English)
This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic.
Curation method
The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books.
Dataset summary
The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.openlibrary_dump_2024-04-30
OpenLibrary Dump (2024-04-30)
This dataset contains the OpenLibrary dump of April 2024 converted to Parquet and DuckDB for easier querying.
Formats
Original GZIP dumps
The original GZIP dumps are available at data/dumps. The dumps are gzipped TSV files with the original OL JSON record contained in the fifth column of the TSV.
DuckDB
The authors, works and editions dumps were imported as tables into… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/openlibrary_dump_2024-04-30.Japanese_Bandori_Band_Story
Japanese Bandori Band Story
Japanese Band Story text retrieved from the Bestdori scenario assets.
This snapshot contains 26 story entries, 493 chapters,
and 30679 rows (28800 dialogue rows).
Created at 2026-09-15T02:11:27.707570+00:00.
Files
data/train-*.parquet: Hub dataset shards generated by Dataset.push_to_hub.
data/band_stories.jsonl: local combined dataset, also included in the downloadable ZIP.
stories/story_XXXX/: complete per-story TXT, CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Japanese_Bandori_Band_Story.storyscope
StoryScope
stories_train.parquet, stories_val.parquet, stories_test.parquet, stories_dev.parquet: prompt metadata plus AI-generated stories from GPT-5.4, Claude Sonnet 4.6, DeepSeek V3.2, Kimi K2.5, and Gemini 3 Flash
storyscope_features.parquet: 304 extracted narrative features for 61,575 story rows
taxonomy.json: the 304-feature taxonomy spanning 10 narrative dimensions
models/: trained XGBoost classifiers for binary human-vs-AI detection and 6-way authorship attribution… See the full description on the dataset page: https://huggingface.co/datasets/jjrussell10/storyscope.story-imprinting
Story Imprinting — training datasets
Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble.
Paper · Code
Contents
Paper section
Folder
Data
3.1 — Sabotage
3_1_sabotage/
Three training mixtures and separate sabotage/clean story pools
3.2 — Narration preferences
3_2_narration_preferences/
Six training mixtures and 12 story pools
4 — Affinity
4_selectivity/
Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.mm-long-storytelling-bench
MM Long Storytelling Bench — v3
⚠️ The 756 model-drafted questions have been WITHDRAWN from this
dataset's splits (2026-08-05). They were drafted by a model that is also
an evaluation target, which makes them circular as a measurement
instrument. They are kept in full, with the reasoning, under
data/v3/archive/ — nothing was deleted.
The splits currently hold 6 worked examples (status: "example"),
which document the required format and are not a benchmark. Do not use
this… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mm-long-storytelling-bench.story_writing_benchmark
Story Evaluation Dataset
This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks.
This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.reddit-story-niche-classification-dataset
🧠 Reddit Niche Classification Dataset
This dataset contains 13,061 Reddit posts annotated with a custom niche label (e.g. advice, drama, humor, unknown, etc). It includes structured features engineered from post metadata, not raw text — making it ideal for lightweight classification models.
🧾 Schema
Column
Type
Description
title
string
Post title
selftext
string
Post body text
subreddit
string
Subreddit the post belongs to
flair
string
Flair… See the full description on the dataset page: https://huggingface.co/datasets/atin5551/reddit-story-niche-classification-dataset.fw-story-shard7Danbooru-Dataset-csv
Danbooru Dataset CSV
面向 Danbooru 标签管理 / 打标工具的公开元数据合集。这里只放整理后的 CSV,不含任何图片。后续还会继续补充 artist、copyright 等更多表;本页只做项目总览,各文件以仓库里的 CSV 为准。
标签与 wiki 来自 Danbooru。本仓库整理表使用 MIT 协议。原图版权仍归各自作者。
当前文件
文件
内容
截止日期
行数
danbooru_dataset_general_260820.csv
general 通用标签(别名、层级、父子、分类、wiki)
2026-08-20
106,414
danbooru_character_tags.csv
character 角色标签(别名、作品、父标签、投稿数)
2026-07-20
329,747
danbooru_artist_tags.csv
artist 画师标签(译名、数据量)
—
576,842
tag-near-synonym-relations4.csv… See the full description on the dataset page: https://huggingface.co/datasets/StoryAura/Danbooru-Dataset-csv.storyworld-plays
Storyworld Plays
This dataset is an append-friendly collection of public, executable
storyworld-evaluation records. Its first shard contains the completed
schema-guided Jinn Town campaign from Jinn or Beast? Theological Identity
Frames as Alignment Surfaces in Small Language Models.
The shard contains:
192 public turns from 24 two-seat, eight-turn episodes;
three development worlds, four fixed constitutional assemblies, and two
paired seeds;
the executed LDT, TRM, and SRT… See the full description on the dataset page: https://huggingface.co/datasets/AlephFunk/storyworld-plays.StorySparkQA
StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children’s Story-Based Learning
This repository contains the StorySparkQA dataset for our paper: StorySparkQA: A Dataset for Narrative Comprehension with External Commonsense Knowledge for Children Education.
The StorySparkQA dataset is constructed based on FairytaleQA, which contains CSV file of 278 fairytale stories from Project Gutenberg and a set of questions and answer pairs (QA-pairs) developed by… See the full description on the dataset page: https://huggingface.co/datasets/NEU-HAI/StorySparkQA.storyweaver-writing-zh
StoryWeaver 中文写作质量评测集
12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。
来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html
核心结论
接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。
k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。
题目怎么设计的
每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.kimi-k3-story-corpus-embeddings
Kimi K3 Story Corpus with Gemini Embeddings
V1 vs. V2: Use V2 for new work. V1 is the original generation built with the legacy Simula prompt taxonomy, where narration/POV and delivery medium were partly combined and second-person or document-shaped stories appeared too often. V2 is a fresh regeneration from revised Simula prompts: grammatical person/focalization and delivery medium are separated, complexification is disabled, the strategy set is simplified, and prompts are… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/kimi-k3-story-corpus-embeddings.merged_simple_story_10k_1story_generation_rl
Story Generation RL (EpisodeBench)
This dataset is the reinforcement-learning (RL) training resource released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL.
EpisodeBench represents each story as an episode graph with explicit states, observable trigger-conditioned transitions, and interaction budgets, turning long-form narrative progression into a measurable evaluation object. The Story Generation RL… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_rl.adaption-pokemon-story-prompts
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-pokemon_story_prompts
This dataset contains prompts instructing a model to write stories about specific Pokémon based on their detailed attributes, including stats, types, abilities, and lore. Each entry provides structured data such as height, weight, generation, and flavor text alongside an image URL. The primary focus is on generating creative narratives grounded… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/adaption-pokemon-story-prompts.story_analogyStoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding
This is the StoryAnalogy dataset in the EMNLP'23 paper: StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding.
If you use this research, please cite us:
@inproceedings{jiayang2023storyanalogy,
title={StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical… See the full description on the dataset page: https://huggingface.co/datasets/JoeyCheng/story_analogy.hathi_pd_books_us_ia_2024-05-01devto-war-story-performance
Dev.to War-Story Performance Dataset
941 articles published on dev.to under the @whoffagents account, spanning April 2026. Includes title, tags, engagement metrics, and reading time.
Why this exists
We run an agentic content pipeline that publishes developer war-stories daily. This dataset captures real performance data across article formats to answer: what titles and tags actually get reactions on dev.to?
Preliminary finding: war-story framing ("I did X and here's what… See the full description on the dataset page: https://huggingface.co/datasets/WH0FF/devto-war-story-performance.balinese-story-texts-extendsWorkflow
tw-kid-story-0.26M
Dataset Card for tw-kid-story-0.26M
本資料集收錄繁體中文兒童故事(童話、寓言、科普故事等)文本,總 token 數約 0.26M(26 萬);可作為兒少教材、親子共讀 chatbot 等應用的繁中模型補強語料。
Dataset Details
Dataset Description
資料以兒童取向之短篇故事為主,文體特色:
句式短、節奏快,適合朗讀。
角色設定鮮明(如:魔法師、勇敢的孩子、善良的動物)。
多有寓意或品格教育的結尾。
token 數以 meta-llama/Llama-3.2-3B 之 tokenizer 計算約 0.26M;樣本數小於 1K,每篇故事為一筆樣本。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese
License: cc-by-nc-sa-4.0
Dataset Sources
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-kid-story-0.26M.arxiv_KL_story_v1_featureshathi_full_20240501gemma-3-27b-it_storygen-prompts-200
google/gemma-3-27b-it — storygen-prompts-200
Model outputs from the micro-creativity inference suite.
Model: google/gemma-3-27b-it
Dataset: storygen-prompts-200 (200 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_storygen-prompts-200.mistral-small-3.2-24b-instruct-2506_storygen-prompts-200
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — storygen-prompts-200
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: storygen-prompts-200 (200 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_storygen-prompts-200.story-classifier-Instruct-v1StorySet
Dataset Card for 123Tapwater/StorySet
Dataset Description
StorySet is a curated subset of manu/project_gutenberg. This dataset contains:
Full text of classic fiction books
5000+ Samples
Cleaned text preserving chapter structure
Genre classifications with confidence scores
Contents
The dataset contains:
5 core fiction genres: Mystery, Romance, Science Fiction, General Fiction, and Children's literature
Cleaned text versions with:
Metadata headers/footers… See the full description on the dataset page: https://huggingface.co/datasets/123Tapwater/StorySet.
