CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tiendung /cnen_novels TODOs Lấy: vi_docln.net cn_qidian.com en_novelhall Làm sạch data Loại bỏ You can read the novel online free at novelhall.com Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank Lựa chọn novels để train theo ranking từ cao tới thấp Notes: _cn_pixiv-novel, _cn_sis-novel bỏ vì có nội dung 18+ _en_webnovel.com bỏ vì nội dung trùng với en_novelhall Crawled… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/cnen_novels.textn<1K0 likes9.5k downloads3y agoHugging Face02tiendung /novels TODOs Lấy: vi_docln.net cn_qidian.com en_novelhall Crawl novel ranking from https://www.webnovel.com/ranking/novel/all_time/popular_rank (done) Lựa chọn novels để train theo ranking từ cao tới thấp => bỏ vì chỉ lấy đc ranking của top 200 truyện Tạm thời lọc theo độ dài text của novels (ưu tiên các novel dài) Và loại bỏ You can read the novel online free at novelhall.com => Tìm kiếm và show mội lines có chứa keywords novelhall.com Notes: _cn_pixiv-novel… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/novels.textn<1K1 likes8.9k downloads3y agoHugging Face03b3x0m /Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier. Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make. (͡ ° ͜ʖ ͡ °) tags: socks, garter belt, foot fetish, ntr, netori..... Thanks Moleys/Numeron for the dataset donation. texttext-classification100M<n<1B247 likes1.5k downloads2y agoHugging Face04OmniAICreator /Japanese-Novels-23Mgated Japanese-Novels-23M This dataset contains Japanese web novels that I collected personally. Machine-Learning Use OnlyAccess is restricted to bona fide machine-learning–related purposes.To request access, please provide a detailed explanation of the specific tasks or applications for which you intend to use the dataset. Total records: 23,212,809 Total characters: 80,846,120,027 Total tokens (Llama 4 tokenizer): 55,406,468,406 (55.4 B) tabulartext-generation10M<n<100M29 likes911 downloads1y agoHugging Face05jetaudio /chinese_web_novelsgatedOver 200.000 Chinese web novels. text100K<n<1M28 likes513 downloads3y agoHugging Face06qhchina /100_top_chinese_novelstext100K<n<1M1 likes442 downloads1y agoHugging Face07roettger /eighteenth_century_french_novels General information This dataset contains 12 Mio Token of Literary French prose 1751-1800 in plain text format, built within the project 'Mining and Modeling Text' (2019-2023) at Trier University. For the dataset in XML/TEI see the GitHub repository of the project. Collection de romans français du dix-huitième siècle (1751-1800) / Collection of Eighteenth-Century French Novels (1751-1800) This collection of Eighteenth-Century French Novels contains 200 digital… See the full description on the dataset page: https://huggingface.co/datasets/roettger/eighteenth_century_french_novels.texttext-generation100K<n<1M2 likes383 downloads2y agoHugging Face08mrzjy /Chinese_interactive_novels_3k 中文互动小说结构化语料 This dataset contains uncleaned (!) 3534 structured Chinese interactive novels (中文互动小说), accounting for around 0.25B (gpt-3.5) tokens in total. All contents are parsed from certain online sources. Usage This dataset can be potentially used for LLM training. But be aware that you'd better clean the data yourself to remove undesired low-quality contents. Each novel is a dict structured as follows: class Novel: book_title: str book_author: str… See the full description on the dataset page: https://huggingface.co/datasets/mrzjy/Chinese_interactive_novels_3k.tabulartext-generation1K<n<10K10 likes313 downloads2y agoHugging Face09alpindale /light-novelstext1M<n<10M42 likes209 downloads3y agoHugging Face10jetaudio /zh_novels_senstext10M<n<100M1 likes169 downloads2y agoHugging Face11hugfaceguy0001 /Novelstext1K<n<10K1 likes109 downloads2y agoHugging Face12sunorme /scholar-novels-curatedscholar-novels-curated 精挑细选+清洗过的小说,目标风格[仙侠,玄幻,武侠] 只用于个人学习和研究大模型预训练用途 texttext-generation10K<n<100K1 likes91 downloads6mo agoHugging Face13atsushi3110 /novels-jatext100K<n<1M6 likes68 downloads2y agoHugging Face14molbal /dramallama-novels DramaLlama dataset This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset. Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune Step 1: Getting novels We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time. I'm running the following scripts: pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/molbal/dramallama-novels.texttext-generation10K<n<100K4 likes57 downloads2y agoHugging Face15Volko76 /french-novels-18th-shuffletext10K<n<100K0 likes50 downloads10mo agoHugging Face16sunorme /Chinese-H-NovelsUpdate 12/07/2024: convert to parquet to download easier. Chinese 18+ novels corpus, use at your own risk, you and only you are responsible for every choice you make. (͡ ° ͜ʖ ͡ °) tags: socks, garter belt, foot fetish, ntr, netori..... Thanks Moleys/Numeron for the dataset donation. texttext-classification100M<n<1B1 likes49 downloads6mo agoHugging Face17UppsalaNLP /swedish-novels-1800-1940 Swedish Novels 1800-1940 A corpus of Swedish literary novels and short story collections from 1800-1940, sourced from Litteraturbanken. Dataset Description This dataset contains over 1.1 million sentences from 350 Swedish literary works spanning 140 years of Swedish literature. The texts have been sentence-segmented and include metadata about authors, titles, and publication years. Fields text: The sentence text author: Author name title: Work title year:… See the full description on the dataset page: https://huggingface.co/datasets/UppsalaNLP/swedish-novels-1800-1940.tabulartext-generation1M<n<10M0 likes37 downloads8mo agoHugging Face18Krixter /Novelstext100K<n<1M0 likes37 downloads4mo agoHugging Face19lanesun /pixiv-novelstext100K<n<1M0 likes36 downloads1y agoHugging Face20winglian /visual-novels-jsontext1K<n<10K10 likes35 downloads3y agoHugging Face21Pclanglais /Brahe-NovelsThe Brahe-Novels dataset is a collection of annotated novel excerpts in the public domain. It was originally created to train Brahe, an LLM fine-tuned for literary analysis. Most of the texts come from the Gutenberg project. The annotations include a mix of synthetic data and manual annotations. In accordance with the principles laid out by the US copyright office, all synthetic data and hybrid synthetic data are in the public domain as well. text1K<n<10K6 likes33 downloads3y agoHugging Face22LizRob6913 /dramallama-novels DramaLlama dataset This is the dataset repository of DramaLlama. This repository contains scripts designed to gather and prepare the dataset. Note: This repository builds upon the findings of https://github.com/molbal/llm-text-completion-finetune Step 1: Getting novels We will use The Gutenberg project again to gather novels. Let's get some drama categories. I will aim for a larger dataset size this time. I'm running the following scripts: pip install requests… See the full description on the dataset page: https://huggingface.co/datasets/LizRob6913/dramallama-novels.texttext-generation10K<n<100K0 likes27 downloads9mo agoHugging Face23PJMixers-Dev /DSULT-Core_light-novels-v2-axotext1K<n<10K0 likes25 downloads10mo agoHugging Face24Matt300209 /NovelSelect_reformattedtext10K<n<100K0 likes23 downloads1y agoHugging Face25Delta-Vector /Ursa-Light-Novels-Largetext100K<n<1M6 likes22 downloads1y agoHugging Face26y1yang0 /scholar-novels-curatedscholar-novels-curated 精挑细选+清洗过的小说,目标风格[仙侠,玄幻,武侠] 只用于个人学习和研究大模型预训练用途 texttext-generationn<1K3 likes20 downloads6mo agoHugging Face27dearprakash /tamil_novels license: cc-by-4.0 This dataset is an attempt to collect the list of work that are in the public domain or work that is shared under Creative Commons license The novels are in Tamil language and the format is plain text format. Name of work Source கடலுக்கு அப்பால்-ப.சிங்காரம் அழிசி புயலிலே ஒரு தோணி - ப.சிங்காரம் அழிசி சத்திய சோதனை-மகாத்மா காந்தி அழிசி நவகாளி யாத்திரை-சாவி அழிசி Work from the following will be added soon படைப்புகள் நாட்டுடைமை ஆக்கப்பட்ட… See the full description on the dataset page: https://huggingface.co/datasets/dearprakash/tamil_novels.text10K<n<100K1 likes19 downloads2y agoHugging Face28NewEden-Forge /Light-Novels-ShareGPTtext1K<n<10K0 likes19 downloads1y agoHugging Face29jetaudio /novels_pro_nerCreated by using Gemini-Pro-2.5 with the following System Prompt: You are a machine-like, high-precision annotator for Chinese web novels. Your only function is to rewrite a given paragraph by embedding specific tags. You must adhere to the following rules with absolute strictness. Mistakes are not acceptable; it is better to leave a word untagged than to tag it incorrectly. **Output Format:** Rewrite the original paragraph completely. Enclose an identified entity in square brackets `[]` and… See the full description on the dataset page: https://huggingface.co/datasets/jetaudio/novels_pro_ner.text100K<n<1M0 likes18 downloads1y agoHugging Face30512duncanl /wh40k_novelsManually cleaned version of Warhammer 40k novels that contains rows of text that are 10000 characters long, with the last 500 characters being an overlap with the first 500 characters in the next row. Does not include contents page (and dramatis personae) or afterword. texttext-generation1K<n<10K1 likes17 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.