CoolFace
18 results

gutenberg

sedthh /gutenberg_english Dataset Card for Project Gutenber - English Language eBooks A collection of non-english language eBooks (48284 rows, 80%+ of all english language books available on the site) from the Project Gutenberg site with metadata removed. Originally colected for https://github.com/LAION-AI/Open-Assistant (follows the OpenAssistant training format) The METADATA column contains catalogue meta information on each book as a serialized JSON: key original column language - text_id… See the full description on the dataset page: https://huggingface.co/datasets/sedthh/gutenberg_english.texttext-generation10K<n<100K39 likes12k downloads4y agoHugging Facemanu /project_gutenberg Dataset Card for "Project Gutenberg" Project Gutenberg is a library of over 70,000 free eBooks, hosted at https://www.gutenberg.org/. All examples correspond to a single book, and contain a header and a footer of a few lines (delimited by a *** Start of *** and *** End of *** tags). Usage from datasets import load_dataset ds = load_dataset("manu/project_gutenberg", split="fr", streaming=True) print(next(iter(ds))) License Full license is available here:… See the full description on the dataset page: https://huggingface.co/datasets/manu/project_gutenberg.texttext-generation10K<n<100K74 likes8.9k downloads3y agoHugging FaceChristophSchuhmann /1-sentence-level-gutenberg-en_arxiv_pubmed_sodatext100M<n<1B1 likes5.4k downloads3y agoHugging Facecommon-pile /project_gutenberg Project Gutenberg Description Project Gutenberg is an online collection of over 75,000 digitized books available as plain text. We use all books that are 1) English and 2) marked as in the Public Domain according to the provided metadata. Additionally, we include any books that are part of the PG19 dataset, which only includes books that are over 100 years old. Minimal preprocessing is applied to remove the Project Gutenberg header and footers, but many scanned books… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/project_gutenberg.texttext-generation10K<n<100K5 likes3.4k downloads1y agoHugging Faceweaverlabs /gutenberg-conversations The Gutenberg Conversations Dataset A comprehensive collection meticulously curated from the extensive library of Project Gutenberg. This dataset specifically focuses on conversational excerpts from a diverse range of literary works, spanning various genres and time periods. It is designed to support and advance research in natural language processing, conversational analysis, machine learning, and linguistics. Each entry in the dataset represents a conversational excerpt… See the full description on the dataset page: https://huggingface.co/datasets/weaverlabs/gutenberg-conversations.text10K<n<100K1 likes3.4k downloads2y agoHugging Faceargilla /gutenberg_spacy-ner Dataset Card for "gutenberg_spacy-ner" More Information needed textn<1K4 likes2.5k downloads3y agoHugging Face