datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cnen_novels
TODOs
Lấy:
vi_docln.net
cn_qidian.com
en_novelhall
Làm sạch data
Loại bỏ You can read the novel online free at novelhall.com
Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank
Lựa chọn novels để train theo ranking từ cao tới thấp
Notes:
_cn_pixiv-novel, _cn_sis-novel bỏ vì có nội dung 18+
_en_webnovel.com bỏ vì nội dung trùng với en_novelhall
Crawled… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/cnen_novels.novels
TODOs
Lấy:
vi_docln.net
cn_qidian.com
en_novelhall
Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank (done)
Lựa chọn novels để train theo ranking từ cao tới thấp => bỏ vì chỉ lấy đc ranking của top 200 truyện
Tạm thời lọc theo độ dài text của novels (ưu tiên các novel dài)
Và loại bỏ You can read the novel online free at novelhall.com
=> Tìm kiếm và show mội lines có chứa keywords novelhall.com
Notes:
_cn_pixiv-novel… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/novels.Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier.
Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make.
(͡ ° ͜ʖ ͡ °)
tags: socks, garter belt, foot fetish, ntr, netori.....
Thanks Moleys/Numeron for the dataset donation.
Japanese-Novels-23M
Japanese-Novels-23M
This dataset contains Japanese web novels that I collected personally.
Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset.
Total records: 23,212,809
Total characters: 80,846,120,027
Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B)
generated-novelschinese_web_novelsOver 200.000 Chinese web novels.
tsinghua_novels_zhtw100_top_chinese_novelseighteenth_century_french_novels
General information
This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University.
For the dataset in XML/TEI see the GitHub repository of the project.
Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800)
This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.visual-novels
Visual Novel Dataset
This dataset contains parsed Visual Novel scripts for training language models. The dataset consists of approximately 60 million tokens of parsed scripts.
Dataset Structure
The dataset follows a general structure for visual novel scripts:
Dialogue lines: Dialogue lines are formatted with the speaker's name followed by a colon, and the dialogue itself enclosed in quotes. For example:
John: "Hello, how are you?"
Actions and narration: Actions and… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/visual-novels.NovelSumThis repository hosts the data accompanying the ACL 2025 main conference paper "Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric".
📋 Overview
In this research, we tackle the fundamental challenge of accurately measuring dataset diversity for instruction tuning and introduce NovelSum, a reliable diversity metric that jointly accounts for inter-sample distances and information density, and shows a strong correlation with model performance.… See the full description on the dataset page: https://huggingface.co/datasets/Sirius518/NovelSum.Chinese_interactive_novels_3k
中文互动小说结构化语料
This dataset contains uncleaned (!) 3534 structured Chinese interactive novels (中文互动小说), accounting for around 0.25B (gpt-3.5) tokens in total.
All contents are parsed from certain online sources.
Usage
This dataset can be potentially used for LLM training. But be aware that you'd better clean the data yourself to remove undesired low-quality contents.
Each novel is a dict structured as follows:
class Novel:
book_title: str
book_author: str… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/Chinese_interactive_novels_3k.NovelStorylight-novelszh_novels_sensnovels-tsinghuavisual-novels-v1.1
Visual Novels v1.1 (WIP)
This dataset contains parsed Visual Novel scripts.
Dataset Structure
We provide 2 variants of the same dataset:
Extracted (Not Yet released)
Contains the raw unformatted version directly extracted from the game's files.
jsonl
Parsed versions of the raw extracted versions. A sample is provided below:
{
"meta": {
"game": "<Game Title>.jsonl",
"scene_key": "<Scene Key>"
},
"namedconversation": [
{… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/visual-novels-v1.1.i-love-reading-pixiv-novels
Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels
Dataset Details
Dataset Description
This is a more or less raw dump of pixiv novel data (13,012,017 documents to be exact.)
Are you the hacker?
I scraped pixiv on the same day of the Kadokawa site issues. I had no clue about the issue surrounding nicolive, etc until I noticed after the scrape was done.
around 8 hours before I started the scrape, the websites(?) went down. Soo...… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels.generated-novels_viewlumro_175_novelsNovelsscholar-novels-curatedscholar-novels-curated
精挑细选+清洗过的小说,目标风格[仙侠,玄幻,武侠]
只用于个人学习和研究大模型预训练用途
old_novelsnovels-jaedan20-assignment1-index-selma-novelsdramallama-novels
DramaLlama dataset
This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset.
Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune
Step 1: Getting novels
We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time.
I'm running the following scripts:
pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/molbal/dramallama-novels.french-novels-18th-shuffleChinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier.
Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make.
(͡ ° ͜ʖ ͡ °)
tags: socks, garter belt, foot fetish, ntr, netori.....
Thanks Moleys/Numeron for the dataset donation.
swedish-novels-1800-1940
Swedish Novels 1800-1940
A corpus of Swedish literary novels and short story collections from 1800-1940, sourced from Litteraturbanken.
Dataset Description
This dataset contains over 1.1 million sentences from 350 Swedish literary works spanning 140 years of Swedish literature. The texts have been sentence-segmented and include metadata about authors, titles, and publication years.
Fields
text: The sentence text
author: Author name
title: Work title
year:… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-novels-1800-1940.Novels
