novel
Datasets
All datasets matching “novel”cnen_novels
TODOs
Lấy:
vi_docln.net
cn_qidian.com
en_novelhall
Làm sạch data
Loại bỏ You can read the novel online free at novelhall.com
Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank
Lựa chọn novels để train theo ranking từ cao tới thấp
Notes:
_cn_pixiv-novel, _cn_sis-novel bỏ vì có nội dung 18+
_en_webnovel.com bỏ vì nội dung trùng với en_novelhall
Crawled… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/cnen_novels.novels
TODOs
Lấy:
vi_docln.net
cn_qidian.com
en_novelhall
Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank (done)
Lựa chọn novels để train theo ranking từ cao tới thấp => bỏ vì chỉ lấy đc ranking của top 200 truyện
Tạm thời lọc theo độ dài text của novels (ưu tiên các novel dài)
Và loại bỏ You can read the novel online free at novelhall.com
=> Tìm kiếm và show mội lines có chứa keywords novelhall.com
Notes:
_cn_pixiv-novel… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/novels.novel-agent-sft-dataset
All Novel Can Be Galgame — 完整数据集
中文小说叙事理解项目的完整数据集。包含 669 本中文小说的原始文本、标注和训练数据,用于训练叙事 Agent 系统。
项目地址:https://github.com/lin1753/novel2galgame
训练代码仓库:https://github.com/lin1753/novel-agent
数据规模
目录
文件数
大小
说明
training/
52
689 MB
训练用 SFT 数据 (JSONL)
raw-books/
671
327 MB
669 本原始小说
processed/
39,842
1.2 GB
按章节预处理文本
annotations/
1,626
1 MB
原始标注文件
合计
42,191
2.2 GB
目录结构
datasets/
├── training/
│ ├── base-sft/… See the full description on the dataset page: https://huggingface.co/datasets/mikuhhn1239/novel-agent-sft-dataset.LLM_for_novel_with_history_and_superheroines
LLM for novel with history and superheroines
This dataset contains the complete text of several novel, fully generated with the assistance of open-source large language models (LLMs). It represents an experiment in large-scale literary creation under a human-AI collaborative paradigm: the human author designed the story architecture, character settings, historical research, and narrative pacing, while open-source LLMs carried out text expansion, dialogue generation, scene… See the full description on the dataset page: https://huggingface.co/datasets/MARK18964758/LLM_for_novel_with_history_and_superheroines.chinese-novel-nonH-collect
Dataset Card for Dataset Name
license: cc0-1.0
task_categories:
- text-classification
- summarization
language:
- zh
tags:
- art
size_categories:
- 100M<n<1B
novelty-bench
NoveltyBench
Prompts and evaluation results for
NoveltyBench, which measures how many of a
model's ten sampled responses to a prompt are meaningfully different, and how
good those are. Code: novelty-bench/novelty-bench.
Prompts
data/curated-*.parquet (100 prompts written for the benchmark) and
data/wildchat-*.parquet (1,000 prompts from real ChatGPT conversations), with
fields id and prompt. These load as the curated and wildchat splits of
the default config.… See the full description on the dataset page: https://huggingface.co/datasets/yimingzhang/novelty-bench.
