datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cnen_novels
TODOs
Lấy:
vi_docln.net
cn_qidian.com
en_novelhall
Làm sạch data
Loại bỏ You can read the novel online free at novelhall.com
Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank
Lựa chọn novels để train theo ranking từ cao tới thấp
Notes:
_cn_pixiv-novel, _cn_sis-novel bỏ vì có nội dung 18+
_en_webnovel.com bỏ vì nội dung trùng với en_novelhall
Crawled… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/cnen_novels.novels
TODOs
Lấy:
vi_docln.net
cn_qidian.com
en_novelhall
Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank (done)
Lựa chọn novels để train theo ranking từ cao tới thấp => bỏ vì chỉ lấy đc ranking của top 200 truyện
Tạm thời lọc theo độ dài text của novels (ưu tiên các novel dài)
Và loại bỏ You can read the novel online free at novelhall.com
=> Tìm kiếm và show mội lines có chứa keywords novelhall.com
Notes:
_cn_pixiv-novel… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/novels.Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier.
Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make.
(͡ ° ͜ʖ ͡ °)
tags: socks, garter belt, foot fetish, ntr, netori.....
Thanks Moleys/Numeron for the dataset donation.
Japanese-Novels-23M
Japanese-Novels-23M
This dataset contains Japanese web novels that I collected personally.
Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset.
Total records: 23,212,809
Total characters: 80,846,120,027
Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B)
chinese_web_novelsOver 200.000 Chinese web novels.
100_top_chinese_novelseighteenth_century_french_novels
General information
This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University.
For the dataset in XML/TEI see the GitHub repository of the project.
Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800)
This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.Chinese_interactive_novels_3k
中文互动小说结构化语料
This dataset contains uncleaned (!) 3534 structured Chinese interactive novels (中文互动小说), accounting for around 0.25B (gpt-3.5) tokens in total.
All contents are parsed from certain online sources.
Usage
This dataset can be potentially used for LLM training. But be aware that you'd better clean the data yourself to remove undesired low-quality contents.
Each novel is a dict structured as follows:
class Novel:
book_title: str
book_author: str… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/Chinese_interactive_novels_3k.light-novelszh_novels_sensNovelsscholar-novels-curatedscholar-novels-curated
精挑细选+清洗过的小说,目标风格[仙侠,玄幻,武侠]
只用于个人学习和研究大模型预训练用途
novels-jadramallama-novels
DramaLlama dataset
This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset.
Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune
Step 1: Getting novels
We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time.
I'm running the following scripts:
pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/molbal/dramallama-novels.french-novels-18th-shuffleChinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier.
Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make.
(͡ ° ͜ʖ ͡ °)
tags: socks, garter belt, foot fetish, ntr, netori.....
Thanks Moleys/Numeron for the dataset donation.
swedish-novels-1800-1940
Swedish Novels 1800-1940
A corpus of Swedish literary novels and short story collections from 1800-1940, sourced from Litteraturbanken.
Dataset Description
This dataset contains over 1.1 million sentences from 350 Swedish literary works spanning 140 years of Swedish literature. The texts have been sentence-segmented and include metadata about authors, titles, and publication years.
Fields
text: The sentence text
author: Author name
title: Work title
year:… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-novels-1800-1940.Novelspixiv-novelsvisual-novels-jsonBrahe-NovelsThe Brahe-Novels dataset is a collection of annotated novel excerpts in the public domain. It was originally created to train Brahe, an LLM fine-tuned for literary analysis.
Most of the texts come from the Gutenberg project.
The annotations include a mix of synthetic data and manual annotations. In accordance with the principles laid out by the US copyright office, all synthetic data and hybrid synthetic data are in the public domain as well.
dramallama-novels
DramaLlama dataset
This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset.
Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune
Step 1: Getting novels
We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time.
I'm running the following scripts:
pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/LizRob6913/dramallama-novels.DSULT-Core_light-novels-v2-axoNovelSelect_reformattedUrsa-Light-Novels-Largescholar-novels-curatedscholar-novels-curated
精挑细选+清洗过的小说,目标风格[仙侠,玄幻,武侠]
只用于个人学习和研究大模型预训练用途
tamil_novels
license: cc-by-4.0
This dataset is an attempt to collect the list of work that are in the public domain or work that is shared under Creative Commons license
The novels are in Tamil language and the format is plain text format.
Name of work
Source
கடலுக்கு அப்பால்-ப.சிங்காரம்
அழிசி
புயலிலே ஒரு தோணி - ப.சிங்காரம்
அழிசி
சத்திய சோதனை-மகாத்மா காந்தி
அழிசி
நவகாளி யாத்திரை-சாவி
அழிசி
Work from the following will be added soon
படைப்புகள் நாட்டுடைமை ஆக்கப்பட்ட… See the full description on the dataset page: https://huggingface.co/datasets/dearprakash/tamil_novels.Light-Novels-ShareGPTnovels_pro_nerCreated by using Gemini-Pro-2.5 with the following System Prompt:
You are a machine-like, high-precision annotator for Chinese web novels. Your only function is to rewrite a given paragraph by embedding specific tags. You must adhere to the following rules with absolute strictness. Mistakes are not acceptable; it is better to leave a word untagged than to tag it incorrectly.
**Output Format:**
Rewrite the original paragraph completely. Enclose an identified entity in square brackets `[]` and… See the full description on the dataset page: https://huggingface.co/datasets/jetaudio/novels_pro_ner.wh40k_novelsManually cleaned version of Warhammer 40k novels that contains rows of text that are 10000 characters long, with the last 500 characters being an overlap with the first 500 characters in the next row. Does not include contents page (and dramatis personae) or afterword.
