CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01truthful-ai /story-imprinting Story Imprinting — training datasets Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble. Paper · Code Contents Paper section Folder Data 3.1 — Sabotage 3_1_sabotage/ Three training mixtures and separate sabotage/clean story pools 3.2 — Narration preferences 3_2_narration_preferences/ Six training mixtures and 12 story pools 4 — Affinity 4_selectivity/ Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.tabulartext-generation100K<n<1M0 likes328 downloads9d agoHugging Face02luoojason /mm-long-storytelling-bench MM Long Storytelling Bench — v3 ⚠️ The 756 model-drafted questions have been WITHDRAWN from this dataset's splits (2026-08-05). They were drafted by a model that is also an evaluation target, which makes them circular as a measurement instrument. They are kept in full, with the reasoning, under data/v3/archive/ — nothing was deleted. The splits currently hold 6 worked examples (status: "example"), which document the required format and are not a benchmark. Do not use this… See the full description on the dataset page: https://huggingface.co/datasets/luoojason/mm-long-storytelling-bench.tabularquestion-answeringn<1K0 likes213 downloads2mo agoHugging Face03PinkPixel /Story-Writing 📖 Story-Writing Dataset ✨ This dataset is a collection of creative writing stories based on the Writing Prompts ([WP]) format. It is designed to help models learn how to write compelling, structured, and emotionally engaging narratives. 📂 Dataset Structure The data is provided in ChatML format, making it ideal for instruction tuning. Files writing_train_chatml.jsonl: Training data. writing_valid_chatml.jsonl: Validation data. Example Entry {… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Story-Writing.texttext-generation1M<n<10M3 likes200 downloads5mo agoHugging Face04nothingiisreal /Short-Storygen-v2I'd recommend you use this dataset instead because its not slopped unlike this one Short Stories generated by Opus. Original dataset by Sao10K text1K<n<10K10 likes119 downloads2y agoHugging Face05AlephFunk /storyworld-plays Storyworld Plays This dataset is an append-friendly collection of public, executable storyworld-evaluation records. Its first shard contains the completed schema-guided Jinn Town campaign from Jinn or Beast? Theological Identity Frames as Alignment Surfaces in Small Language Models. The shard contains: 192 public turns from 24 two-seat, eight-turn episodes; three development worlds, four fixed constitutional assemblies, and two paired seeds; the executed LDT, TRM, and SRT… See the full description on the dataset page: https://huggingface.co/datasets/AlephFunk/storyworld-plays.tabular1K<n<10K0 likes104 downloads2mo agoHugging Face06schonsense /human_ai_story_contrastive_v4 Human–AI Story Contrastive Six-Source Dataset Dataset summary This dataset contains 1,443 matched short-story prompt groups with human and machine-written realizations drawn from six source populations. The dataset is organized one row per group_id rather than one row per text. Each row preserves the common writing prompt, the human reference story, the available earlier machine generations, the starting-policy generation, and two GPT-5.6 Sol fields added for… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v4.text1K<n<10K0 likes96 downloads9d agoHugging Face07lesserfield /fgo-storytexttranslation10K<n<100K3 likes77 downloads2y agoHugging Face08PinkPixel /Childrens-Story-Writing 🧒 Children's Story Writing Dataset ✨ This dataset is a collection of creative short stories written for children. It is designed to help models learn child-friendly language and how to follow specific narrative instructions (e.g., incorporating specific features or sentences). 📂 Dataset Structure The data is provided in ChatML format, making it ideal for instruction tuning. Files writing_train_children.jsonl: Training data. writing_valid_children.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/PinkPixel/Childrens-Story-Writing.texttext-generation1M<n<10M2 likes77 downloads5mo agoHugging Face09rx1lora /StoryPlay_RolePlay-NPCv2 RolePlay-NPCv2 The newest RP dataset containing some high-quality dataset for Gemma3NPC. We combined pippa, NPC-Dialogue_v2, Sonnet-Roleplay and ReLe_Synthetic_v1_json. WARNING -- Some conversations contain highly NSFW content, use it with caution! texttext-generation10K<n<100K1 likes77 downloads2mo agoHugging Face10schonsense /human_ai_story_contrastive_v2Collapsed and Augmented with llama3.3sft on-policy/policy adjacent responses Prepare for some AI generated descriptions. Dataset Construction and Statistics Purpose: A contrastive creative-writing dataset for studying distributional stylistic differences between human-written and model-generated text. Total size: 5,733 texts grouped across 1,443 unique prompts / human responses. Source breakdown: 1,443 human responses 1,421 GPT-3.5 responses 1,426 Claude Opus responses 1,443… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v2.text1K<n<10K0 likes77 downloads20d agoHugging Face11godwei123 /storyweaver-writing-zh StoryWeaver 中文写作质量评测集 12 道按写作失效模式反推设计的中文创作题、4 个参赛者写出的 48 篇章节、432 条逐维度两两判决(含裁判完整推理原文)。 来自 StoryWeaver 的写作质量评测轨道。榜单:https://storyweaver.cn/benchmark-writing.html 核心结论 接系统比换一代底模更管用。同一底模接上多 Agent 系统后的胜率:k2.5 **75.1%**、k2.6 **60.2%**;而 k2.5(系统) 对 k2.6(裸) 是 70.3%,反过来只有 37.2%——系统加持能把旧一代底模抬过裸的新一代底模。系统档拿下 22 个维度里的 20 个榜首,包括全部 9 个负向维度。 k2.5 与 k2.6 之间 54.7%,落在噪音带内,不构成结论。 题目怎么设计的 每道题咬住 rubric 里的一个维度或负向维度,用硬约束逼出功力:… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-writing-zh.tabulartext-generationn<1K1 likes76 downloads2mo agoHugging Face12Dans-DiscountModels /RUCAIBox-Story-Generation-Alpacahttps://huggingface.co/datasets/RUCAIBox/Story-Generation RUC AI Box HC Story Generation augmented and converted to alpaca format. No filtering has been done. texttext-generation1K<n<10K13 likes65 downloads3y agoHugging Face13NeuralNovel /Neural-Story-v1 Neural-Story-v1 Dataset Overview The Neural-Story-v1 dataset is a curated collection of short stories featuring a rich variety of genres and plot settings. Carefully assembled by NeuralNovel, this dataset aims to serve as a valuable resource for testing and fine-tuning small language models using LoRa. Data Source The dataset content is a result of a combination of automated generation by Mixtral 8x7b and manual refinement. Purpose Designed… See the full description on the dataset page: https://huggingface.co/datasets/NeuralNovel/Neural-Story-v1.textn<1K10 likes58 downloads3y agoHugging Face14k007-a /story-summary Short Story Summarization Dataset Short stories from agentlans/euclaise-WritingPromptsX Summarized using Qwen/Qwen3.5-9B with the following prompt: Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/k007-a/story-summary.texttext-generation10K<n<100K0 likes58 downloads1mo agoHugging Face15nassimjp /pashto-reasoning-children-story-crafting-dataset Pashto Reasoning Children Story Crafting Dataset Welcome to the Pashto Reasoning Children Story Crafting Dataset! This dataset is designed to empower Large Language Models (LLMs) with the capability to craft engaging, moral, and logically structured children's stories in the Pashto language, integrating explicit reasoning steps. Dataset Overview & Methodology Language: Pashto (ps) Base Prompts: 100 unique core story prompts. Total Samples: 500 diverse story… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-reasoning-children-story-crafting-dataset.texttext-generationn<1K0 likes54 downloads6d agoHugging Face16schonsense /human_ai_story_contrastive_v3Regenerated and Cleaned of off task responses. py .\audit_rebuilt_pi0_shortstory_v2.py ` >> --old .\human_ai_story_pi0\augmented_groups.jsonl ` >> --new .\human_ai_story_pi0_shortstory_v2\augmented_groups.jsonl ` >> --out-dir .\human_ai_story_pi0_shortstory_v2\rebuild_audit { "status": "PASS", "rows_old": 1443, "rows_new": 1443, "unique_old": 1443, "unique_new": 1443, "integrity": { "source_prompt_mutations": 0, "human_response_mutations": 0… See the full description on the dataset page: https://huggingface.co/datasets/schonsense/human_ai_story_contrastive_v3.text1K<n<10K0 likes47 downloads11d agoHugging Face17h34v7 /story-sessionstextn<1K0 likes40 downloads1y agoHugging Face18schonsense /human_ai_story_contrastive_v2a_audit text1K<n<10K0 likes40 downloads19d agoHugging Face19PJMixers /Gryphe_Opus-WritingPrompts-Story2Prompt-ShareGPTtext1K<n<10K1 likes39 downloads2y agoHugging Face20Severian /Internal-Knowledge-Map-StoryWriter-RolePlaying This is an Expansion and Subset of the Internal Knowledge Map dataset that focuses on Story Writing and Role Playing. I was curious to see if I could adapt my IKM structure and approach to improve Story Telling, Role Playing/Character/Discourse in an LLM. Here are 2,071 highly-detailed and unique examples that allow an LLM to exhibit more depth, diverse perspectives and novel interactions. Side benefit is the LLM also writes in well-formed, aesthetically pleasing formatting and is an… See the full description on the dataset page: https://huggingface.co/datasets/Severian/Internal-Knowledge-Map-StoryWriter-RolePlaying.text1K<n<10K15 likes36 downloads3y agoHugging Face21sarahooker /adaption-pokemon-story-prompts This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. adaption-pokemon_story_prompts This dataset contains prompts instructing a model to write stories about specific Pokémon based on their detailed attributes, including stats, types, abilities, and lore. Each entry provides structured data such as height, weight, generation, and flavor text alongside an image URL. The primary focus is on generating creative narratives grounded… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/adaption-pokemon-story-prompts.image1K<n<10K0 likes36 downloads5mo agoHugging Face22build-small-hackathon /bedtime-story-machine-trace 🌙 Bedtime Story Machine — Agent Build Trace This dataset contains the build trace for the Bedtime Story Machine project, built for the Build Small Hackathon. What's in this trace Complete build steps from concept to deployment Architecture decisions and model choices Challenges encountered and solutions Tools and infrastructure used Project → Live App → GitHub textn<1K0 likes35 downloads4mo agoHugging Face23AtlasUnified /atlas-storytellertext1K<n<10K9 likes34 downloads3y agoHugging Face24Threatthriver /Hindi-story-news Hindi Web Content Dataset Overview This dataset contains a collection of Hindi text data scraped from various websites. The data was collected using a domain-restricted scraper that extracts paragraphs of text from specified domains. The dataset includes content from news articles, literature, and other web pages. The scraped text has been stored in JSON format and is intended for use in natural language processing (NLP) tasks, such as language modeling, text generation… See the full description on the dataset page: https://huggingface.co/datasets/Threatthriver/Hindi-story-news.text1K<n<10K2 likes34 downloads2y agoHugging Face25agentlans /story-summary Short Story Summarization Dataset Short stories from agentlans/euclaise-WritingPromptsX Summarized using Qwen/Qwen3.5-9B with the following prompt: Summarize the short story below in a single well-written paragraph that captures the main plot, key characters, central conflict, and resolution while preserving the original meaning and tone. Focus only on the most important details, avoid unnecessary specifics or minor subplots, and do not add interpretations or information that… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/story-summary.texttext-generation10K<n<100K0 likes34 downloads3mo agoHugging Face26Babelscape /story-summeval Dataset Card for Story-SummEval Dataset Description For a thorough description of the data creation please refer to the ACL 2024 paper: "FENICE: Factuality Evaluation of summarization based on NLI and Claim Extraction", Scirè et al. (2024). Summary This dataset contains summaries of stories from Gutenberg and Wikisource along with their factuality labels. Summaries are generated from several models provided by the paper "Echoes from Alexandria" by Scirè et al.… See the full description on the dataset page: https://huggingface.co/datasets/Babelscape/story-summeval.textn<1K8 likes31 downloads2y agoHugging Face27godwei123 /storyweaver-chunking-zh StoryWeaver 中文叙事切分语料 12 篇中文叙事短文,按切分算法的失效模式反推设计,用于比较 chunking 策略在长篇小说创作场景下的表现。 来自 StoryWeaver 的切分质量评测轨道。榜单:https://storyweaver.cn/benchmark-chunking.html 这批语料的特别之处 它不是真实小说的随机采样,而是每篇专门写来触发某一类切分失效: 对话密集、闪回嵌套、隐晦转场、双线交替、同场景多视角、同场景话题漂移…… 这么设计是为了让不同算法拉开差距——真实文本里这些情况稀疏出现,抽样评测容易得到"各方法差不多"的钝化结论。代价是分布有偏,解读结果时必须计入这一点。 文件 corpus.jsonl(12 行) 字段 说明 doc_id 篇名(中文,如 夜班、面馆) text 全文 chars 字数 leaderboard.json 首轮评测的汇总结果:6… See the full description on the dataset page: https://huggingface.co/datasets/godwei123/storyweaver-chunking-zh.texttext-generationn<1K0 likes29 downloads2mo agoHugging Face28virtualkevin /tell-me-a-story Tell Me A Story This dataset mirrors the Tell Me A Story dataset from DeepMind's google-deepmind/tell_me_a_story repository. The source repository is associated with the paper Agents' Room: Narrative Generation through Multi-step Collaboration. Source The source GitHub repository was cloned from: google-deepmind/tell_me_a_story Source commit used for this upload: e4910ea1d2bae82efcaf8ba9fde50ab3a419320e The upstream repository points to encrypted JSONL files hosted in… See the full description on the dataset page: https://huggingface.co/datasets/virtualkevin/tell-me-a-story.textn<1K1 likes24 downloads4mo agoHugging Face29ImagineIt /Updated_story-datasettext1K<n<10K1 likes23 downloads3y agoHugging Face30lianghsun /tw-kid-story-0.26Mgated Dataset Card for tw-kid-story-0.26M 本資料集收錄繁體中文兒童故事(童話、寓言、科普故事等)文本,總 token 數約 0.26M(26 萬);可作為兒少教材、親子共讀 chatbot 等應用的繁中模型補強語料。 Dataset Details Dataset Description 資料以兒童取向之短篇故事為主,文體特色: 句式短、節奏快,適合朗讀。 角色設定鮮明(如:魔法師、勇敢的孩子、善良的動物)。 多有寓意或品格教育的結尾。 token 數以 meta-llama/Llama-3.2-3B 之 tokenizer 計算約 0.26M;樣本數小於 1K,每篇故事為一筆樣本。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional Chinese License: cc-by-nc-sa-4.0 Dataset Sources Repository:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-kid-story-0.26M.tabulartext-generationn<1K0 likes22 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.