CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mikuhhn1239 /novel-agent-sft-dataset All Novel Can Be Galgame — 完整数据集 中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。 项目地址:https://github.com/lin1753/novel2galgame 训练代码仓库:https://github.com/lin1753/novel-agent 数据规模 目录 文件数 大小 说明 training/ 52 689 MB 训练用 SFT 数据 (JSONL) raw-books/ 671 327 MB 669 本原始小说 processed/ 39,842 1.2 GB 按章节预处理文本 annotations/ 1,626 1 MB 原始标注文件 合计 42,191 2.2 GB 目录结构 datasets/ ├── training/ │ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.texttext-generation10K<n<100K6 likes4.8k downloads3mo agoHugging Face02Bryan35406 /fable-novel-eightsday fable: eightsday 한국어판 제목: 「fable — 여드레날」 License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md. The story of a man who realized the world is one enormous language model. 세상이 하나의 거대한 언어 모델임을 깨달은 남자의 이야기. A complete Korean–English bilingual serialized novel and a section-aligned literary parallel corpus, co-written by a human… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/fable-novel-eightsday.imagetext-generationn<1K1 likes244 downloads1mo agoHugging Face03LooksJuicy /Chinese-Roleplay-Novel一直以来,中文角色扮演开源数据集更关注超拟人方向或纯角色对话方向,严重缺乏交互游戏方向的开源数据,因此许多模型尤其参数量较小的模型对酒馆类的角色卡支持较差。 为了解决这一困境,本项目抛砖引玉,基于4500条小说文本使用GPT4o构建出约260条酒馆style的数据集,均为多轮对话,每轮对话都包括状态数据,如时间、角色状态、任务进度等。 数据key对应含义如下: world:表示当前故事的世界观,通常可以加入到system prompt中 scence:表示当前故事发生场景,包括时间、地点、环境、任务目标 character:表示当前故事中可能出现的角色和对应简介 field:表示这条数据每轮对话中需要生成的状态信息 conversations:表示这条数据的对话内容,分为问候语、主角(user)和系统(assistant) fields_format:表示状态信息的填充格式prompt,可能是列表、表格、JSON等各种形式 format_list:表示状态信息的填充结果 状态信息的示例如下 **健康状态**: 🌿 良好,身体颤抖 **精神状态**: 🌟 恐惧,极度紧张… See the full description on the dataset page: https://huggingface.co/datasets/LooksJuicy/Chinese-Roleplay-Novel.textn<1K88 likes186 downloads2y agoHugging Face04jetaudio /zh_novels_senstext10M<n<100M1 likes172 downloads2y agoHugging Face05yuanhezhang /lean4-stat-learning-theory-novel A Large-Scale Lean 4 Dataset on Statistical Learning Theory We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-novel.texttext-generationn<1K0 likes158 downloads8mo agoHugging Face06taozi555 /novel-rpgated Novel-RP: Multilingual Novel Role-Playing Dataset A multilingual novel-based role-playing dataset for training and evaluating LLMs on character persona simulation. 📖 Overview Novel-RP is a multilingual role-playing dataset built from web novels and role-playing conversations, specifically designed for training large language models on character role-playing tasks. This dataset contains two main subsets: train: Novel-based role-playing data (ShareGPT format) - from… See the full description on the dataset page: https://huggingface.co/datasets/taozi555/novel-rp.tabulartext-generation10K<n<100K11 likes151 downloads9mo agoHugging Face07miyuki2026 /chinese_porn_noveltabular1M<n<10M6 likes150 downloads8mo agoHugging Face08Crownelius /Novel-Crafting-250x Novel-Crafting-250x Stats Metric Value Total prompt tokens 157,900 Total completion tokens 1,021,978 Total tokens 1,179,878 Total cost $5.27 (USD) Average turns 1.00 Average tool calls 0.00 Average tokens per row 2,359.76 Cost estimated using Unknown pricing on OpenRouter ($1.0/M input, $5.0/M output) textn<1K5 likes147 downloads2mo agoHugging Face09asd567557275 /chinese_novel Space Grimoire Novel Corpus (Traditional Chinese) Full text of the Traditional Chinese web novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage), split by chapter. Field Value Author 睡半夜怎麼三更 License CC BY 4.0 Language Traditional Chinese (zh-Hant-TW) Genre Fantasy, steampunk, political intrigue Chapters 283 (interludes included) Parts 10 Paragraphs 30,760 Characters (body text) 1,737,806 Version 2026-09-13 Companion dataset… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/chinese_novel.tabulartext-generationn<1K3 likes145 downloads11d agoHugging Face10v3ucn /chinese-novel-datasettextn<1K5 likes128 downloads2y agoHugging Face11taozi555 /novel_texttexttext-generation100K<n<1M3 likes126 downloads2y agoHugging Face12atsushi3110 /novels-jatext100K<n<1M6 likes119 downloads2y agoHugging Face13silk-road /ChatHaruhi_NovelWritingtext10K<n<100K8 likes118 downloads3y agoHugging Face14Dxniz /Novelist Dataset Card for Novelist Dataset Summary Novelist is a synthetic creative-writing and narrative-reasoning dataset designed for long-context fiction systems, scene planners, continuity-aware story models, multilingual literary translation, and child-safe TinyStories generation. The dataset mixes direct prose, explicit reasoning traces, quality-only judge outputs, full-book artifacts, and multilingual translation outputs inside a single narrative training ecosystem. This… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Novelist.tabulartext-generation10K<n<100K2 likes117 downloads6mo agoHugging Face15NewEden-Forge /Light-Novels-ShareGPTtext1K<n<10K0 likes109 downloads1y agoHugging Face16Orion-zhen /tagged-pixiv-noveltext10K<n<100K10 likes96 downloads2y agoHugging Face17PericlesSavio /novel17_testtext1K<n<10K1 likes95 downloads3y agoHugging Face18silk-road /50-Chinese-Novel-Characterstext100K<n<1M12 likes88 downloads3y agoHugging Face19fsAX /chinese-novel-datasettextn<1K0 likes80 downloads2y agoHugging Face20vikasaivyas /hindi-novel-sft-dataset 📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस) यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है। 🌟 प्रमुख विशेषताएँ (Key Highlights) 10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ। 100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.texttext-generationn<1K0 likes68 downloads15d agoHugging Face21nineoneone /chinese-h-noveltext100K<n<1M8 likes60 downloads2y agoHugging Face22Seikaijyu /Sex-novel-filtered 色情小说数据集 本数据集包含了3392条单条数据最大长度2500token的数据集 这是一个被人工精细化清洗过的色情小说数据集,此数据来源于Pixiv小说板块 原数据集有3w条,我花了一个通宵的时间配合正则人工清洗了它,最终得到了3000条语料 虽然精细处理过,但不能保证百分百干净 虽然这么说.....但此数据已经可以直接训练了,至少不会有什么大问题 另外提一嘴,现代网络小说真难练啊,ctx特长,质量特低,风格逻辑混乱,收敛特慢,感觉根本就是一无是处嘛 text1K<n<10K93 likes55 downloads2y agoHugging Face23jas0nhuan9 /chinese-noveltextn<1K0 likes55 downloads2y agoHugging Face24Orion-zhen /pixiv-noveltext10K<n<100K25 likes53 downloads3y agoHugging Face25werty1248 /Korean-1930-Novel-Scene-Summarize 한국 저작권 만료 소설에 대한 씬 분리 및 요약 데이터 셋 원천 데이터 출처: https://gongu.copyright.or.kr/gongu/wrt/wrtCl/listWrtText.do?menuNo=200019 총 96개 소설 수집 및 전처리 한자가 많은 소설 제외 한자 제거, 띄어쓰기 전처리 수행 씬 분리 사용 모델: Gemini-1.5-Flash (띄어쓰기 포함) 100자 이상, 1200자 미만으로 적절한 문장에서 씬 단위로 분리하도록 지시 총 12,108씬 생성 요약 사용 모델: Gemini-1.5-Flash(때때로 GPT-4o) 각 Scene에서 인물, 주요 소품, 사건을 추출하고, 요약(scenario)을 생성하도록 함 textsummarization10K<n<100K4 likes51 downloads2y agoHugging Face26Minami-su /Anime_novel_datasetstexttext-generation10K<n<100K37 likes49 downloads3y agoHugging Face27Dxniz /novelist-cot-writing-raw-v1 Novelist: Human-Like Creative Writing Dataset (RAW) This dataset is designed to train LLMs in high-quality creative writing. It focuses on narrative depth, coherent world-building, and logical character psychology. The data was generated using DeepSeek-R1. Dataset Overview We focused on Quality over Quantity. The goal was to move away from generic "AI slop" and create text that feels grounded and intentional. Total Tokens: ~29.4 Million Total Examples: 3,369 Format:… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/novelist-cot-writing-raw-v1.tabular1K<n<10K1 likes49 downloads7mo agoHugging Face28open-llm-leaderboard /Alsebay__Qwen2.5-7B-test-novelist-detailsgated Dataset Card for Evaluation run of Alsebay/Qwen2.5-7B-test-novelist Dataset automatically created during the evaluation run of model Alsebay/Qwen2.5-7B-test-novelist The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Alsebay__Qwen2.5-7B-test-novelist-details.tabular10K<n<100K0 likes48 downloads2y agoHugging Face29sms25444 /sex-noveltext1K<n<10K3 likes47 downloads2y agoHugging Face30zacll /chinese-adult-novel-v0.2text10K<n<100K11 likes46 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.