CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adams-story /datacomp200m Datacomp200m This is a smaller version of the datacomp_1b dataset. Filtering was done by taking all rows that had self similarity (inner product) above 0.32. This resulted in 213009083 (213 million) rows. The results of the datacomp paper suggest that filtering by CLIP score is better than random sampling. Included in this repo are search indices created using autofaiss, over the text and image embeddings. There are two ways to access metadata, either in .parquet files in the… See the full description on the dataset page: https://huggingface.co/datasets/adams-story/datacomp200m.image100M<n<1B3 likes110k downloads3y agoHugging Face02storytracer /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.tabulartext-generation100K<n<1M191 likes4.5k downloads3y agoHugging Face03storytracer /LoC-PD-Books Library of Congress Public Domain Books (English) This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic. Curation method The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books. Dataset summary The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.tabulartext-generation10K<n<100K42 likes1.4k downloads3y agoHugging Face04storytracer /openlibrary_dump_2024-04-30 OpenLibrary Dump (2024-04-30) This dataset contains the OpenLibrary dump of April 2024 converted to Parquet and DuckDB for easier querying. Formats Original GZIP dumps The original GZIP dumps are available at data/dumps. The dumps are gzipped TSV files with the original OL JSON record contained in the fifth column of the TSV. DuckDB The authors, works and editions dumps were imported as tables into… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/openlibrary_dump_2024-04-30.tabular10M<n<100M0 likes1k downloads2y agoHugging Face05KomeijiForce /Japanese_Bandori_Band_Story Japanese Bandori Band Story Japanese Band Story text retrieved from the Bestdori scenario assets. This snapshot contains 26 story entries, 493 chapters, and 30679 rows (28800 dialogue rows). Created at 2026-09-15T02:11:27.707570+00:00. Files data/train-*.parquet: Hub dataset shards generated by Dataset.push_to_hub. data/band_stories.jsonl: local combined dataset, also included in the downloadable ZIP. stories/story_XXXX/: complete per-story TXT, CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Japanese_Bandori_Band_Story.tabulartext-generation10K<n<100K0 likes604 downloads12d agoHugging Face06jjrussell10 /storyscope StoryScope stories_train.parquet, stories_val.parquet, stories_test.parquet, stories_dev.parquet: prompt metadata plus AI-generated stories from GPT-5.4, Claude Sonnet 4.6, DeepSeek V3.2, Kimi K2.5, and Gemini 3 Flash storyscope_features.parquet: 304 extracted narrative features for 61,575 story rows taxonomy.json: the 304-feature taxonomy spanning 10 narrative dimensions models/: trained XGBoost classifiers for binary human-vs-AI detection and 6-way authorship attribution… See the full description on the dataset page: https://huggingface.co/datasets/jjrussell10/storyscope.tabulartext-classification10K<n<100K6 likes372 downloads6mo agoHugging Face07truthful-ai /story-imprinting Story Imprinting — training datasets Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble. Paper · Code Contents Paper section Folder Data 3.1 — Sabotage 3_1_sabotage/ Three training mixtures and separate sabotage/clean story pools 3.2 — Narration preferences 3_2_narration_preferences/ Six training mixtures and 12 story pools 4 — Affinity 4_selectivity/ Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.tabulartext-generation100K<n<1M0 likes328 downloads9d agoHugging Face08luoojason /mm-long-storytelling-bench MM Long Storytelling Bench — v3 ⚠️ The 756 model-drafted questions have been WITHDRAWN from this dataset's splits (2026-08-05). They were drafted by a model that is also an evaluation target, which makes them circular as a measurement instrument. They are kept in full, with the reasoning, under data/v3/archive/ — nothing was deleted. The splits currently hold 6 worked examples (status: "example"), which document the required format and are not a benchmark. Do not use this… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mm-long-storytelling-bench.tabularquestion-answeringn<1K0 likes213 downloads2mo agoHugging Face09lars1234 /story_writing_benchmark Story Evaluation Dataset This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks. This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.tabulartext-generation10K<n<100K6 likes204 downloads2y agoHugging Face10atin5551 /reddit-story-niche-classification-dataset 🧠 Reddit Niche Classification Dataset This dataset contains 13,061 Reddit posts annotated with a custom niche label (e.g. advice, drama, humor, unknown, etc). It includes structured features engineered from post metadata, not raw text — making it ideal for lightweight classification models. 🧾 Schema Column Type Description title string Post title selftext string Post body text subreddit string Subreddit the post belongs to flair string Flair… See the full description on the dataset page: https://huggingface.co/datasets/atin5551/reddit-story-niche-classification-dataset.tabulartext-classification10K<n<100K1 likes185 downloads1y agoHugging Face11liyucheng /fw-story-shard7tabular1M<n<10M0 likes173 downloads2y agoHugging Face12StoryAura /Danbooru-Dataset-csv Danbooru Dataset CSV 面向 Danbooru 标签管理 / 打标工具的公开元数据合集。这里只放整理后的 CSV,不含任何图片。后续还会继续补充 artist、copyright 等更多表;本页只做项目总览,各文件以仓库里的 CSV 为准。 标签与 wiki 来自 Danbooru。本仓库整理表使用 MIT 协议。原图版权仍归各自作者。 当前文件 文件 内容 截止日期 行数 danbooru_dataset_general_260820.csv general 通用标签(别名、层级、父子、分类、wiki) 2026-08-20 106,414 danbooru_character_tags.csv character 角色标签(别名、作品、父标签、投稿数) 2026-07-20 329,747 danbooru_artist_tags.csv artist 画师标签(译名、数据量) — 576,842 tag-near-synonym-relations4.csv… See the full description on the dataset page: https://huggingface.co/datasets/StoryAura/Danbooru-Dataset-csv.tabulartext-classification100K<n<1M1 likes143 downloads10d agoHugging Face13AlephFunk /storyworld-plays Storyworld Plays This dataset is an append-friendly collection of public, executable storyworld-evaluation records. Its first shard contains the completed schema-guided Jinn Town campaign from Jinn or Beast? Theological Identity Frames as Alignment Surfaces in Small Language Models. The shard contains: 192 public turns from 24 two-seat, eight-turn episodes; three development worlds, four fixed constitutional assemblies, and two paired seeds; the executed LDT, TRM, and SRT… See the full description on the dataset page: https://huggingface.co/datasets/AlephFunk/storyworld-plays.tabular1K<n<10K0 likes104 downloads2mo agoHugging Face14NEU-HAI /StorySparkQA StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children’s Story-Based Learning This repository contains the StorySparkQA dataset for our paper: StorySparkQA: A Dataset for Narrative Comprehension with External Commonsense Knowledge for Children Education. The StorySparkQA dataset is constructed based on FairytaleQA, which contains CSV file of 278 fairytale stories from Project Gutenberg and a set of questions and answer pairs (QA-pairs) developed by… See the full description on the dataset page: https://huggingface.co/datasets/NEU-HAI/StorySparkQA.tabularquestion-answering1K<n<10K2 likes83 downloads2y agoHugging Face15godwei123 /storyweaver-writing-zh StoryWeaver 中文写作质量评测集 12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。 来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html 核心结论 接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。 k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。 题目怎么设计的 每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.tabulartext-generationn<1K1 likes76 downloads2mo agoHugging Face16Mercity /kimi-k3-story-corpus-embeddings Kimi K3 Story Corpus with Gemini Embeddings V1 vs. V2: Use V2 for new work. V1 is the original generation built with the legacy Simula prompt taxonomy, where narration/POV and delivery medium were partly combined and second-person or document-shaped stories appeared too often. V2 is a fresh regeneration from revised Simula prompts: grammatical person/focalization and delivery medium are separated, complexification is disabled, the strategy set is simplified, and prompts are… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/kimi-k3-story-corpus-embeddings.documenttext-generation1K<n<10K2 likes58 downloads1mo agoHugging Face17sghosts /merged_simple_story_10k_1tabular100K<n<1M0 likes57 downloads11mo agoHugging Face18HeAAAAA /story_generation_rl Story Generation RL (EpisodeBench) This dataset is the reinforcement-learning (RL) training resource released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. EpisodeBench represents each story as an episode graph with explicit states, observable trigger-conditioned transitions, and interaction budgets, turning long-form narrative progression into a measurable evaluation object. The Story Generation RL… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_rl.tabulartext-generation10K<n<100K1 likes40 downloads2mo agoHugging Face19sarahooker /adaption-pokemon-story-prompts This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-pokemon_story_prompts This dataset contains prompts instructing a model to write stories about specific Pokémon based on their detailed attributes, including stats, types, abilities, and lore. Each entry provides structured data such as height, weight, generation, and flavor text alongside an image URL. The primary focus is on generating creative narratives grounded… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/adaption-pokemon-story-prompts.image1K<n<10K0 likes36 downloads5mo agoHugging Face20JoeyCheng /story_analogyStoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding This is the StoryAnalogy dataset in the EMNLP'23 paper: StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding. If you use this research, please cite us: @inproceedings{jiayang2023storyanalogy, title={StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical… See the full description on the dataset page: https://huggingface.co/datasets/JoeyCheng/story_analogy.tabular10K<n<100K1 likes33 downloads2y agoHugging Face21storytracer /hathi_pd_books_us_ia_2024-05-01tabular100K<n<1M0 likes31 downloads2y agoHugging Face22WH0FF /devto-war-story-performance Dev.to War-Story Performance Dataset 941 articles published on dev.to under the @whoffagents account, spanning April 2026. Includes title, tags, engagement metrics, and reading time. Why this exists We run an agentic content pipeline that publishes developer war-stories daily. This dataset captures real performance data across article formats to answer: what titles and tags actually get reactions on dev.to? Preliminary finding: war-story framing ("I did X and here's what… See the full description on the dataset page: https://huggingface.co/datasets/WH0FF/devto-war-story-performance.tabulartext-classificationn<1K0 likes26 downloads5mo agoHugging Face23pradanaadn /balinese-story-texts-extendsWorkflow tabulartext-classification1K<n<10K1 likes24 downloads11mo agoHugging Face24lianghsun /tw-kid-story-0.26Mgated Dataset Card for tw-kid-story-0.26M 本資料集收錄繁體中文兒童故事(童話、寓言、科普故事等)文本,總 token 數約 0.26M(26 萬);可作為兒少教材、親子共讀 chatbot 等應用的繁中模型補強語料。 Dataset Details Dataset Description 資料以兒童取向之短篇故事為主,文體特色: 句式短、節奏快,適合朗讀。 角色設定鮮明(如:魔法師、勇敢的孩子、善良的動物)。 多有寓意或品格教育的結尾。 token 數以 meta-llama/Llama-3.2-3B 之 tokenizer 計算約 0.26M;樣本數小於 1K,每篇故事為一筆樣本。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional Chinese License: cc-by-nc-sa-4.0 Dataset Sources Repository:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-kid-story-0.26M.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face25mehuldamani /arxiv_KL_story_v1_featurestabularn<1K0 likes21 downloads4mo agoHugging Face26storytracer /hathi_full_20240501tabular10M<n<100M0 likes20 downloads2y agoHugging Face27ZachW /gemma-3-27b-it_storygen-prompts-200 google/gemma-3-27b-it — storygen-prompts-200 Model outputs from the micro-creativity inference suite. Model: google/gemma-3-27b-it Dataset: storygen-prompts-200 (200 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt sent to the model (after meta-prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/gemma-3-27b-it_storygen-prompts-200.tabulartext-generationn<1K0 likes17 downloads5mo agoHugging Face28ZachW /mistral-small-3.2-24b-instruct-2506_storygen-prompts-200 mistralai/Mistral-Small-3.2-24B-Instruct-2506 — storygen-prompts-200 Model outputs from the micro-creativity inference suite. Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506 Dataset: storygen-prompts-200 (200 items) Part of collection: ZachW/llm-creativity-benchmarks Generation config temperature: 0.0 max_tokens: 16384 seed: 42 backend: vllm Columns Column Description task_id Unique task identifier input The exact prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_storygen-prompts-200.tabulartext-generationn<1K0 likes17 downloads5mo agoHugging Face29mehuldamani /story-classifier-Instruct-v1tabular10K<n<100K0 likes16 downloads7mo agoHugging Face30123Tapwater /StorySet Dataset Card for 123Tapwater/StorySet Dataset Description StorySet is a curated subset of manu/project_gutenberg. This dataset contains: Full text of classic fiction books 5000+ Samples Cleaned text preserving chapter structure Genre classifications with confidence scores Contents The dataset contains: 5 core fiction genres: Mystery, Romance, Science Fiction, General Fiction, and Children's literature Cleaned text versions with: Metadata headers/footers… See the full description on the dataset page: https://huggingface.co/datasets/123Tapwater/StorySet.tabulartext-generation1K<n<10K1 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.