datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
novel-agent-sft-dataset
All Novel Can Be Galgame — 完整数据集
中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。
项目地址:https://github.com/lin1753/novel2galgame
训练代码仓库:https://github.com/lin1753/novel-agent
数据规模
目录
文件数
大小
说明
training/
52
689 MB
训练用 SFT 数据 (JSONL)
raw-books/
671
327 MB
669 本原始小说
processed/
39,842
1.2 GB
按章节预处理文本
annotations/
1,626
1 MB
原始标注文件
合计
42,191
2.2 GB
目录结构
datasets/
├── training/
│ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.fable-novel-eightsday
fable: eightsday
한국어판 제목: 「fable — 여드레날」
License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md.
The story of a man who realized the world is one enormous language model.
세상이 하나의 거대한 언어 모델임을 깨달은 남자의 이야기.
A complete Korean–English bilingual serialized novel and a section-aligned literary parallel corpus, co-written by a human… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/fable-novel-eightsday.Chinese-Roleplay-Novel一直以来,中文角色扮演开源数据集更关注超拟人方向或纯角色对话方向,严重缺乏交互游戏方向的开源数据,因此许多模型尤其参数量较小的模型对酒馆类的角色卡支持较差。
为了解决这一困境,本项目抛砖引玉,基于4500条小说文本使用GPT4o构建出约260条酒馆style的数据集,均为多轮对话,每轮对话都包括状态数据,如时间、角色状态、任务进度等。
数据key对应含义如下:
world:表示当前故事的世界观,通常可以加入到system prompt中
scence:表示当前故事发生场景,包括时间、地点、环境、任务目标
character:表示当前故事中可能出现的角色和对应简介
field:表示这条数据每轮对话中需要生成的状态信息
conversations:表示这条数据的对话内容,分为问候语、主角(user)和系统(assistant)
fields_format:表示状态信息的填充格式prompt,可能是列表、表格、JSON等各种形式
format_list:表示状态信息的填充结果
状态信息的示例如下
**健康状态**: 🌿 良好,身体颤抖
**精神状态**: 🌟 恐惧,极度紧张… See the full description on the dataset page: https://huggingface.co/datasets/LooksJuicy/Chinese-Roleplay-Novel.zh_novels_senslean4-stat-learning-theory-novel
A Large-Scale Lean 4 Dataset on Statistical Learning Theory
We present a high-quality, human-verified, large-scale Lean 4 dataset, extracted from our formalization of Statistical Learning Theory (SLT). We present the first comprehensive Lean 4 formalization of SLT grounded in empirical process theory. Our end-to-end formal infrastructure implement the missing contents in latest Lean 4 Mathlib library, including a complete development of Gaussian Lipschitz concentration… See the full description on the dataset page: https://huggingface.co/datasets/yuanhezhang/lean4-stat-learning-theory-novel.novel-rp
Novel-RP: Multilingual Novel Role-Playing Dataset
A multilingual novel-based role-playing dataset for training and evaluating LLMs on character persona simulation.
📖 Overview
Novel-RP is a multilingual role-playing dataset built from web novels and role-playing conversations, specifically designed for training large language models on character role-playing tasks.
This dataset contains two main subsets:
train: Novel-based role-playing data (ShareGPT format) - from… See the full description on the dataset page: https://huggingface.co/datasets/taozi555/novel-rp.chinese_porn_novelNovel-Crafting-250x
Novel-Crafting-250x
Stats
Metric
Value
Total prompt tokens
157,900
Total completion tokens
1,021,978
Total tokens
1,179,878
Total cost
$5.27 (USD)
Average turns
1.00
Average tool calls
0.00
Average tokens per row
2,359.76
Cost estimated using Unknown pricing on OpenRouter ($1.0/M input, $5.0/M output)
chinese_novel
Space Grimoire Novel Corpus (Traditional Chinese)
Full text of the Traditional Chinese web novel 空間魔導書與少年魔法師 (The Space Grimoire and the Young Mage), split by chapter.
Field
Value
Author
睡半夜怎麼三更
License
CC BY 4.0
Language
Traditional Chinese (zh-Hant-TW)
Genre
Fantasy, steampunk, political intrigue
Chapters
283 (interludes included)
Parts
10
Paragraphs
30,760
Characters (body text)
1,737,806
Version
2026-09-13
Companion dataset… See the full description on the dataset page: https://huggingface.co/datasets/asd567557275/chinese_novel.chinese-novel-datasetnovel_textnovels-jaChatHaruhi_NovelWritingNovelist
Dataset Card for Novelist
Dataset Summary
Novelist is a synthetic creative-writing and narrative-reasoning dataset designed for long-context fiction systems, scene planners, continuity-aware story models, multilingual literary translation, and child-safe TinyStories generation. The dataset mixes direct prose, explicit reasoning traces, quality-only judge outputs, full-book artifacts, and multilingual translation outputs inside a single narrative training ecosystem.
This… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Novelist.Light-Novels-ShareGPTtagged-pixiv-novelnovel17_test50-Chinese-Novel-Characterschinese-novel-datasethindi-novel-sft-dataset
📚 Modern Hindi Literature SFT Dataset (आधुनिक हिंदी कथा-साहित्य कॉर्पस)
यह समकालीन आधुनिक हिंदी कथा-साहित्य का सुपरवाइज्ड फाइन-ट्यूनिंग (SFT) डेटासेट है। इसे विशेष रूप से Gemma-2, Llama-3, Mistral आदि मॉडलों को उच्च-कोटि का हिंदी उपन्यास व कहानी लेखन सिखाने के लिए तैयार किया गया है।
🌟 प्रमुख विशेषताएँ (Key Highlights)
10 प्रसिद्ध आधुनिक पुस्तकें: सत्य व्यास, दिव्य प्रकाश दुबे, नीलोत्पल मृणाल एवं नवीन चौधरी की सर्वश्रेष्ठ कृतियाँ।
100% प्रामाणिक मूल पाठ (Zero AI… See the full description on the dataset page: https://huggingface.co/datasets/vikasaivyas/hindi-novel-sft-dataset.chinese-h-novelSex-novel-filtered
色情小说数据集
本数据集包含了3392条单条数据最大长度2500token的数据集
这是一个被人工精细化清洗过的色情小说数据集,此数据来源于Pixiv小说板块
原数据集有3w条,我花了一个通宵的时间配合正则人工清洗了它,最终得到了3000条语料
虽然精细处理过,但不能保证百分百干净
虽然这么说.....但此数据已经可以直接训练了,至少不会有什么大问题
另外提一嘴,现代网络小说真难练啊,ctx特长,质量特低,风格逻辑混乱,收敛特慢,感觉根本就是一无是处嘛
chinese-novelpixiv-novelKorean-1930-Novel-Scene-Summarize
한국 저작권 만료 소설에 대한 씬 분리 및 요약 데이터 셋
원천 데이터 출처: https://gongu.copyright.or.kr/gongu/wrt/wrtCl/listWrtText.do?menuNo=200019
총 96개 소설 수집 및 전처리
한자가 많은 소설 제외
한자 제거, 띄어쓰기 전처리 수행
씬 분리
사용 모델: Gemini-1.5-Flash
(띄어쓰기 포함) 100자 이상, 1200자 미만으로 적절한 문장에서 씬 단위로 분리하도록 지시
총 12,108씬 생성
요약
사용 모델: Gemini-1.5-Flash(때때로 GPT-4o)
각 Scene에서 인물, 주요 소품, 사건을 추출하고, 요약(scenario)을 생성하도록 함
Anime_novel_datasetsnovelist-cot-writing-raw-v1
Novelist: Human-Like Creative Writing Dataset (RAW)
This dataset is designed to train LLMs in high-quality creative writing. It focuses on narrative depth, coherent world-building, and logical character psychology.
The data was generated using DeepSeek-R1.
Dataset Overview
We focused on Quality over Quantity. The goal was to move away from generic "AI slop" and create text that feels grounded and intentional.
Total Tokens: ~29.4 Million
Total Examples: 3,369
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/novelist-cot-writing-raw-v1.Alsebay__Qwen2.5-7B-test-novelist-details
Dataset Card for Evaluation run of Alsebay/Qwen2.5-7B-test-novelist
Dataset automatically created during the evaluation run of model Alsebay/Qwen2.5-7B-test-novelist
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Alsebay__Qwen2.5-7B-test-novelist-details.sex-novelchinese-adult-novel-v0.2
