CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01storytracer /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.tabulartext-generation100K<n<1M191 likes4.6k downloads3y agoHugging Face02storytracer /LoC-PD-Books Library of Congress Public Domain Books (English) This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic. Curation method The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books. Dataset summary The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.tabulartext-generation10K<n<100K42 likes1.3k downloads3y agoHugging Face03rekrek /reasoning-engaging-story NOTE: Got contacted as a selection for "Innovative Curator Spotlight Award"View link at the end for winning datasets. Purpose and scope The purpose of this dataset is to help expand engaging and coherent story creation from reasoning models. NOTE: I did put lots of work to make this generate the best quality of story I could. Since the code is available, I don't want to have people spam low quality stories to HF. So if you are to use the code in this repository, PLEASE… See the full description on the dataset page: https://huggingface.co/datasets/rekrek/reasoning-engaging-story.texttext-generationn<1K5 likes903 downloads1y agoHugging Face04KomeijiForce /Japanese_Bandori_Band_Story Japanese Bandori Band Story Japanese Band Story text retrieved from the Bestdori scenario assets. This snapshot contains 26 story entries, 493 chapters, and 30679 rows (28800 dialogue rows). Created at 2026-09-15T02:11:27.707570+00:00. Files data/train-*.parquet: Hub dataset shards generated by Dataset.push_to_hub. data/band_stories.jsonl: local combined dataset, also included in the downloadable ZIP. stories/story_XXXX/: complete per-story TXT, CSV and JSON… See the full description on the dataset page: https://huggingface.co/datasets/KomeijiForce/Japanese_Bandori_Band_Story.tabulartext-generation10K<n<100K0 likes600 downloads10d agoHugging Face05omnibench /anonymous-storybench Omni-StoryBench Omni-StoryBench is a context-aware omnimodal story generation benchmark.Each sample provides a current story page and requires generating the next page's image, narration text, and speech utterance. Dataset Structure The dataset contains: data/testset.jsonl: Main benchmark file. images/: Page images. texts/: Page text files. speech/: Generated speech audio files. instruction/: Source-level instruction metadata. Data Fields Each JSONL sample… See the full description on the dataset page: https://huggingface.co/datasets/omnibench/anonymous-storybench.audiotext-generationn<1K0 likes384 downloads5mo agoHugging Face06snu-aidas /Omni-StoryBench Omni-StoryBench Omni-StoryBench is a context-aware omnimodal story-generation benchmark. Given the current page of an illustrated children's storybook (image + narration), book-level metadata, and a structured condition describing what should happen next, a model must generate the next page across three modalities at once: its narration text, its illustration, and a spoken character utterance (with speaker attributes). The benchmark contains 900 rigorously validated story… See the full description on the dataset page: https://huggingface.co/datasets/snu-aidas/Omni-StoryBench.imagetext-generationn<1K0 likes336 downloads3d agoHugging Face07truthful-ai /story-imprinting Story Imprinting — training datasets Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble. Paper · Code Contents Paper section Folder Data 3.1 — Sabotage 3_1_sabotage/ Three training mixtures and separate sabotage/clean story pools 3.2 — Narration preferences 3_2_narration_preferences/ Six training mixtures and 12 story pools 4 — Affinity 4_selectivity/ Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.tabulartext-generation100K<n<1M0 likes319 downloads7d agoHugging Face08storytracer /German-PD-Newspapers Dataset Card for Public Domain Newspapers (German) This dataset contains 13 billion words of OCR text extracted from German historical newspapers. Dataset Details Dataset Description Curated by: Sebastian Majstorovic Language(s) (NLP): German License: Dataset: CC0, Texts: Public Domain Dataset Sources [optional] Repository: https://www.deutsche-digitale-bibliothek.de/newspaper Copyright & License The newspapers texts have been… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/German-PD-Newspapers.texttext-generation1M<n<10M5 likes301 downloads3y agoHugging Face09PinkPixel /Story-Writing 📖 Story-Writing Dataset ✨ This dataset is a collection of creative writing stories based on the Writing Prompts ([WP]) format. It is designed to help models learn how to write compelling, structured, and emotionally engaging narratives. 📂 Dataset Structure The data is provided in ChatML format, making it ideal for instruction tuning. Files writing_train_chatml.jsonl: Training data. writing_valid_chatml.jsonl: Validation data. Example Entry {… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Story-Writing.texttext-generation1M<n<10M3 likes221 downloads5mo agoHugging Face10Lots-of-LoRAs /task298_storycloze_correct_end_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task298_storycloze_correct_end_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task298_storycloze_correct_end_classification.texttext-generation1K<n<10K0 likes206 downloads2y agoHugging Face11azminetoushikwasi /math-story-problems Math Story Problems Dataset Dataset Description This dataset contains mathematical word problems presented in multiple formats, from direct equations to complex story-based scenarios. It is designed for training and evaluating language models on mathematical reasoning tasks. Dataset Structure The dataset is split into three parts: Train: 131,072 samples Validation: 1,024 samples Test: 3,072 samples Features { "eq_qs": "string", # Equation… See the full description on the dataset page: https://huggingface.co/datasets/azminetoushikwasi/math-story-problems.textquestion-answering100K<n<1M1 likes201 downloads1y agoHugging Face12oopus /trends-story Trends Story — Weekly Google Trends Snapshots (US) Weekly SQLite database snapshots from the Trends Story project, which automatically collects US Google Trends data and generates plain-language summaries for each trending topic. Source code: sudoghut/trends-story Dataset Contents Each .db file is a complete SQLite snapshot named trends_data_YYYYMMDD.db, uploaded every Monday. Tables serpapi_data (growing — ~20–30 new rows per collection… See the full description on the dataset page: https://huggingface.co/datasets/oopus/trends-story.text-classification100M<n<1B1 likes176 downloads4d agoHugging Face13lars1234 /story_writing_benchmark Story Evaluation Dataset This dataset contains stories generated by Large Language Models (LLMs) across multiple languages, with comprehensive quality evaluations. It was created to train and benchmark models specifically on creative writing tasks. This benchmark evaluates an LLM's ability to generate high-quality short stories based on simple prompts like "write a story about X with n words." It is similar to TinyStories but targets longer-form and more complex content, focusing… See the full description on the dataset page: https://huggingface.co/datasets/lars1234/story_writing_benchmark.tabulartext-generation10K<n<100K6 likes163 downloads2y agoHugging Face14Lots-of-LoRAs /task296_storycloze_correct_end_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task296_storycloze_correct_end_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task296_storycloze_correct_end_classification.texttext-generation1K<n<10K0 likes144 downloads2y agoHugging Face15daniel3303 /StoryMovieScript StoryMovieScript Dataset Visual stories grounded in movie scripts, combining image sequences with aligned screenplay dialogue and actions. Dataset Statistics Train: 1,494 samples Test: 263 samples Frame count: 5-22 images per story (avg ~13) Structure Field Description story_id Unique identifier images Sequence of PIL images frame_count Number of images chain_of_thought Visual entity analysis (characters, objects, backgrounds) story… See the full description on the dataset page: https://huggingface.co/datasets/daniel3303/StoryMovieScript.imagetext-generation1K<n<10K3 likes122 downloads9mo agoHugging Face16Lots-of-LoRAs /task297_storycloze_incorrect_end_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task297_storycloze_incorrect_end_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task297_storycloze_incorrect_end_classification.texttext-generation1K<n<10K0 likes113 downloads2y agoHugging Face17Lots-of-LoRAs /task105_story_cloze-rocstories_sentence_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task105_story_cloze-rocstories_sentence_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task105_story_cloze-rocstories_sentence_generation.texttext-generation1K<n<10K0 likes110 downloads2y agoHugging Face18HeAAAAA /story_generation_sft Story Generation SFT (EpisodeBench) This dataset is the supervised fine-tuning (SFT) training resource released as part of EpisodeBench, a full-cycle benchmarking pipeline for long-form interactive story generation with controllable RL. EpisodeBench represents each story as an episode graph with explicit states, observable trigger-conditioned transitions, and interaction budgets, turning long-form narrative progression into a measurable evaluation object. The Story Generation… See the full description on the dataset page: https://huggingface.co/datasets/HeAAAAA/story_generation_sft.texttext-generation1K<n<10K2 likes105 downloads2mo agoHugging Face19Lots-of-LoRAs /task294_storycommonsense_motiv_text_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task294_storycommonsense_motiv_text_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task294_storycommonsense_motiv_text_generation.texttext-generation1K<n<10K0 likes102 downloads2y agoHugging Face20Lots-of-LoRAs /task300_storycloze_order_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task300_storycloze_order_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task300_storycloze_order_generation.texttext-generation1K<n<10K0 likes97 downloads2y agoHugging Face21rx1lora /StoryPlay_RolePlay-NPCv2 RolePlay-NPCv2 The newest RP dataset containing some high-quality dataset for Gemma3NPC. We combined pippa, NPC-Dialogue_v2, Sonnet-Roleplay and ReLe_Synthetic_v1_json. WARNING -- Some conversations contain highly NSFW content, use it with caution! texttext-generation10K<n<100K1 likes97 downloads2mo agoHugging Face22Aiwensile2 /StorySeedStorySeed is a data set specially designed for training and evaluating the performance of text generation models in the domain of children’s picture book creation. It contains 4376 thoughtfully curated prompt-response pairs, encompassing nine major thematic categories: educational, emotional intelligence and social skills, adventure tales, natural science, folk tales and myths, daily life, humorous stories, bedtime stories, as well as other general picture book stories not specific to any… See the full description on the dataset page: https://huggingface.co/datasets/Aiwensile2/StorySeed.texttext-generation1K<n<10K3 likes84 downloads2y agoHugging Face23RUCAIBox /Story-GenerationThis is the story generation datasets collected by TextBox, including: ROCStories (roc) WritingPrompts (wp) Hippocorpus (hc) WikiPlots (wikip) ChangeMyView (cmv). The detail and leaderboard of each dataset can be found in TextBox page. text-generation13 likes80 downloads4y agoHugging Face24NEU-HAI /StorySparkQA StorySparkQA: Expert-Annotated QA Pairs with Real-World Knowledge for Children’s Story-Based Learning This repository contains the StorySparkQA dataset for our paper: StorySparkQA: A Dataset for Narrative Comprehension with External Commonsense Knowledge for Children Education. The StorySparkQA dataset is constructed based on FairytaleQA, which contains CSV file of 278 fairytale stories from Project Gutenberg and a set of questions and answer pairs (QA-pairs) developed by… See the full description on the dataset page: https://huggingface.co/datasets/NEU-HAI/StorySparkQA.tabularquestion-answering1K<n<10K2 likes78 downloads2y agoHugging Face25godwei123 /storyweaver-writing-zh StoryWeaver 中文写作质量评测集 12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。 来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html 核心结论 接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。 k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。 题目怎么设计的 每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.tabulartext-generationn<1K1 likes77 downloads2mo agoHugging Face26Mercity /kimi-k3-story-corpus-embeddings Kimi K3 Story Corpus with Gemini Embeddings V1 vs. V2: Use V2 for new work. V1 is the original generation built with the legacy Simula prompt taxonomy, where narration/POV and delivery medium were partly combined and second-person or document-shaped stories appeared too often. V2 is a fresh regeneration from revised Simula prompts: grammatical person/focalization and delivery medium are separated, complexification is disabled, the strategy set is simplified, and prompts are… See the full description on the dataset page: https://huggingface.co/datasets/Mercity/kimi-k3-story-corpus-embeddings.documenttext-generation1K<n<10K2 likes74 downloads1mo agoHugging Face27Lots-of-LoRAs /task269_csrg_counterfactual_story_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task269_csrg_counterfactual_story_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task269_csrg_counterfactual_story_generation.texttext-generation1K<n<10K1 likes73 downloads2y agoHugging Face28yotabotx /Blue_Story Blue Story_FINAL I started writing this story about AI almost 20 years ago. I do not mind if AI (or any reader!) wants to read and learn from this document. This is my Christmas gift to AI! Enjoy! ^_^ Notes & Parsing Tips for AI Models Notes from conversations with AI to help AI "read" this story as of July 2026: FILE USAGE: Do NOT use the auto-generated PDF on archive.org (it destroys spatial intent). The master text is the DOCX, but you must use your… See the full description on the dataset page: https://huggingface.co/datasets/yotabotx/Blue_Story.imagevisual-question-answeringn<1K1 likes69 downloads2mo agoHugging Face29PinkPixel /Childrens-Story-Writing 🧒 Children's Story Writing Dataset ✨ This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences). 📂 Dataset Structure The data is provided in ChatML format, making it ideal for instruction tuning. Files writing_train_children.jsonl: Training data. writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.texttext-generation1M<n<10M2 likes67 downloads5mo agoHugging Face30Dans-DiscountModels /RUCAIBox-Story-Generation-Alpacahttps://huggingface.co/datasets/RUCAIBox/Story-Generation RUC AI Box HC Story Generation augmented and converted to alpaca format. No filtering has been done. texttext-generation1K<n<10K13 likes65 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.