CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tiendung /cnen_novels TODOs Lấy: vi_docln.net cn_qidian.com en_novelhall Làm sạch data Loại bỏ You can read the novel online free at novelhall.com Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank Lựa chọn novels để train theo ranking từ cao tới thấp Notes: _cn_pixiv-novel, _cn_sis-novel bỏ vì có nội dung 18+ _en_webnovel.com bỏ vì nội dung trùng với en_novelhall Crawled… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/cnen_novels.textn<1K0 likes9.5k downloads3y agoHugging Face02tiendung /novels TODOs Lấy: vi_docln.net cn_qidian.com en_novelhall Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank (done) Lựa chọn novels để train theo ranking từ cao tới thấp => bỏ vì chỉ lấy đc ranking của top 200 truyện Tạm thời lọc theo độ dài text của novels (ưu tiên các novel dài) Và loại bỏ You can read the novel online free at novelhall.com => Tìm kiếm và show mội lines có chứa keywords novelhall.com Notes: _cn_pixiv-novel… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/novels.textn<1K1 likes8.9k downloads3y agoHugging Face03b3x0m /Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier. Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make. (͡ ° ͜ʖ ͡ °) tags: socks, garter belt, foot fetish, ntr, netori..... Thanks Moleys/Numeron for the dataset donation. texttext-classification100M<n<1B247 likes1.5k downloads2y agoHugging Face04OmniAICreator /Japanese-Novels-23Mgated Japanese-Novels-23M This dataset contains Japanese web novels that I collected personally. Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset. Total records: 23,212,809 Total characters: 80,846,120,027 Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B) tabulartext-generation10M<n<100M29 likes911 downloads1y agoHugging Face05GOAT-AI /generated-novelsdocumentn<1K8 likes572 downloads3y agoHugging Face06jetaudio /chinese_web_novelsgatedOver 200.000 Chinese web novels. text100K<n<1M28 likes513 downloads3y agoHugging Face07aiotvincent /tsinghua_novels_zhtw0 likes485 downloads2mo agoHugging Face08qhchina /100_top_chinese_novelstext100K<n<1M1 likes442 downloads1y agoHugging Face09roettger /eighteenth_century_french_novels General information This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University. For the dataset in XML/TEI see the GitHub repository of the project. Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800) This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.texttext-generation100K<n<1M2 likes383 downloads2y agoHugging Face10alpindale /visual-novels Visual Novel Dataset This dataset contains parsed Visual Novel scripts for training language models. The dataset consists of approximately 60 million tokens of parsed scripts. Dataset Structure The dataset follows a general structure for visual novel scripts: Dialogue lines: Dialogue lines are formatted with the speaker's name followed by a colon, and the dialogue itself enclosed in quotes. For example: John: "Hello, how are you?" Actions and narration: Actions and… See the full description on the dataset page: https://huggingface.co/datasets/alpindale/visual-novels.text-generation69 likes371 downloads3y agoHugging Face11Sirius518 /NovelSumThis repository hosts the data accompanying the ACL 2025 main conference paper "Measuring Data Diversity for Instruction Tuning: A Systematic Analysis and A Reliable Metric". 📋 Overview In this research, we tackle the fundamental challenge of accurately measuring dataset diversity for instruction tuning and introduce NovelSum, a reliable diversity metric that jointly accounts for inter-sample distances and information density, and shows a strong correlation with model performance.… See the full description on the dataset page: https://huggingface.co/datasets/Sirius518/NovelSum.2 likes357 downloads6mo agoHugging Face12mrzjy /Chinese_interactive_novels_3k 中文互动小说结构化语料 This dataset contains uncleaned (!) 3534 structured Chinese interactive novels (中文互动小说), accounting for around 0.25B (gpt-3.5) tokens in total. All contents are parsed from certain online sources. Usage This dataset can be potentially used for LLM training. But be aware that you'd better clean the data yourself to remove undesired low-quality contents. Each novel is a dict structured as follows: class Novel: book_title: str book_author: str… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/Chinese_interactive_novels_3k.tabulartext-generation1K<n<10K10 likes313 downloads2y agoHugging Face13NovelStory /NovelStory2 likes261 downloads11mo agoHugging Face14alpindale /light-novelstext1M<n<10M42 likes209 downloads3y agoHugging Face15jetaudio /zh_novels_senstext10M<n<100M1 likes169 downloads2y agoHugging Face16clemenschen /novels-tsinghua1 likes147 downloads2y agoHugging Face17KaraKaraWitch /visual-novels-v1.1 Visual Novels v1.1 (WIP) This dataset contains parsed Visual Novel scripts. Dataset Structure We provide 2 variants of the same dataset: Extracted (Not Yet released) Contains the raw unformatted version directly extracted from the game's files. jsonl Parsed versions of the raw extracted versions. A sample is provided below: { "meta": { "game": "<Game Title>.jsonl", "scene_key": "<Scene Key>" }, "namedconversation": [ {… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/visual-novels-v1.1.text-generation1 likes140 downloads7d agoHugging Face18DSULT-Core /i-love-reading-pixiv-novels Dataset Card for ilovehentai9000/i-love-reading-pixiv-novels Dataset Details Dataset Description This is a more or less raw dump of pixiv novel data (13,012,017 documents to be exact.) Are you the hacker? I scraped pixiv on the same day of the Kadokawa site issues. I had no clue about the issue surrounding nicolive, etc until I noticed after the scrape was done. around 8 hours before I started the scrape, the websites(?) went down. Soo...… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/i-love-reading-pixiv-novels.text-generation10M<n<100M8 likes121 downloads2y agoHugging Face19chris745820 /generated-novels_viewdocumentn<1K0 likes113 downloads9mo agoHugging Face20upb-nlp /lumro_175_novels1 likes112 downloads1y agoHugging Face21hugfaceguy0001 /Novelstext1K<n<10K1 likes109 downloads2y agoHugging Face22sunorme /scholar-novels-curatedscholar-novels-curated 精挑细选+清洗过的小说,目标风格[仙侠,玄幻,武侠] 只用于个人学习和研究大模型预训练用途 texttext-generation10K<n<100K1 likes91 downloads6mo agoHugging Face23clemenschen /old_novels0 likes71 downloads2y agoHugging Face24atsushi3110 /novels-jatext100K<n<1M6 likes68 downloads2y agoHugging Face25exawken /edan20-assignment1-index-selma-novels0 likes59 downloads7d agoHugging Face26molbal /dramallama-novels DramaLlama dataset This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset. Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune Step 1: Getting novels We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time. I'm running the following scripts: pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/molbal/dramallama-novels.texttext-generation10K<n<100K4 likes57 downloads2y agoHugging Face27Volko76 /french-novels-18th-shuffletext10K<n<100K0 likes50 downloads10mo agoHugging Face28sunorme /Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier. Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make. (͡ ° ͜ʖ ͡ °) tags: socks, garter belt, foot fetish, ntr, netori..... Thanks Moleys/Numeron for the dataset donation. texttext-classification100M<n<1B1 likes49 downloads6mo agoHugging Face29UppsalaNLP /swedish-novels-1800-1940 Swedish Novels 1800-1940 A corpus of Swedish literary novels and short story collections from 1800-1940, sourced from Litteraturbanken. Dataset Description This dataset contains over 1.1 million sentences from 350 Swedish literary works spanning 140 years of Swedish literature. The texts have been sentence-segmented and include metadata about authors, titles, and publication years. Fields text: The sentence text author: Author name title: Work title year:… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-novels-1800-1940.tabulartext-generation1M<n<10M0 likes37 downloads8mo agoHugging Face30Krixter /Novelstext100K<n<1M0 likes37 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.