CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /pre_1929_books_filtered Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.texttext-generation100K<n<1M2 likes1.8k downloads1y agoHugging Face02common-pile /pre_1929_books Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books.texttext-generation100K<n<1M8 likes1.1k downloads1y agoHugging Face03ubaada /booksum-complete-cleaned Description: This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization . This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.textsummarization1K<n<10K23 likes455 downloads2y agoHugging Face04mmichall /BookSummary Polish BookSummary Polish BookSummary is a long-document factual-consistency classification dataset constructed from a publicly available collection of English-language book plot summaries. Each example consists of a short claim followed by a Polish plot summary. The task is a binary classification problem in which the model must determine whether the claim is factually consistent with the information contained in the document. Polish BookSummary is part of the LongContext… See the full description on the dataset page: https://huggingface.co/datasets/mmichall/BookSummary.texttext-classification1K<n<10K0 likes238 downloads22d agoHugging Face05Govi-Rare-Books-Archive /Comprehensive-Antiquarian-and-Rare-Books-Archive Comprehensive Antiquarian & Rare Books Archive Dataset Description This dataset contains pristine, commerce-free bibliographical metadata extracted from the Govi Rare Books Archive. It is engineered to provide high-fidelity, structured historical data for Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines. By supplying ground-truth bibliographical metadata, this repository aims to reduce AI hallucinations and improve semantic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Govi-Rare-Books-Archive/Comprehensive-Antiquarian-and-Rare-Books-Archive.textn<1K0 likes227 downloads19h agoHugging Face06JesseParvess /book_snippets_asrtextn<1K0 likes129 downloads5y agoHugging Face07Blackroot /Tiny-Open-Domain-BooksA tiny example dataset consisting of four books dedicated to the open domain in JSONL format: Alice in Wonderland - Lewis Caroll Dracula - Bram Stoker The Wonderful Wizard of Oz - L. Frank Baum The Count of Monte Cristo - Alexandre Dumas & Auguste Maquet All works are open domain, thus this dataset is also dedicated to the open domain. The dataset has been made to have extremely long context lengths, ideally as close to 2048 at possible without cutting off chunks in strange places. Each… See the full description on the dataset page: https://huggingface.co/datasets/Blackroot/Tiny-Open-Domain-Books.textn<1K5 likes120 downloads3y agoHugging Face08next-token /clean-PD-16000-books3 📚 clean-PD-16000-books3 A treasure trove of ~16,000 high-quality, public domain books in English language — nicely cleaned, with rich metadata, and ready for language modeling. ✨ What Makes This Dataset Special? This isn’t just another dump of dusty old text files. clean-PD-16000-books3 is the result of a rigorous cleaning and curation process applied to a large collection of public domain literature, including: ✅ Readable prose — paragraphized prose, without unnatural… See the full description on the dataset page: https://huggingface.co/datasets/next-token/clean-PD-16000-books3.text10K<n<100K6 likes101 downloads1y agoHugging Face09IsmaelMousa /books Books The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books. This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling. Details The dataset contains 4 columns: title: The tilte of the book. author: The author of the book. category: The… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/books.texttext-generationn<1K5 likes99 downloads2y agoHugging Face10zongqing0068 /EvolvingWorld-Books-Gutenberg Dataset Card for EvolvingWorld-Books-Gutenberg Dataset Summary This dataset provides the 57 source book texts used for EvolvingWorld data construction. The books were selected from Goodreads' Best Books Ever list and obtained from Project Gutenberg. Dataset Structure Each row is a JSON object with the following fields: title: the title of the book author: the author of the book content: the full text content used for data construction Example: {… See the full description on the dataset page: https://huggingface.co/datasets/zongqing0068/EvolvingWorld-Books-Gutenberg.textn<1K1 likes98 downloads3mo agoHugging Face11yuyijiong /Book_Summary_Chinese 中文图书总结数据集 每个样本包含: 图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。 texttext-generationn<1K27 likes86 downloads3y agoHugging Face12freococo /tipitaka_myanmar_translation_books Myanmar Tipitaka Translation (60 Books) This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga. The texts have been converted into a clean, structured JSONL format, suitable for: Natural Language Processing (NLP) LLM Training & Fine-tuning Digital Humanities Research Dhamma Study Applications 📊 Dataset Statistics Total Books: 60 Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.tabulartext-generation100K<n<1M0 likes45 downloads8mo agoHugging Face13mstyslavity /zer0-books scraped_at_utc: 2026-03-24T12:45:50.800626+00:00, source: https://www.collectiveinkbooks.com/zer0-books/our-books/all-books/&s=0, listing_total_books: 383, listing_total_pages: 16, scraped_books: 383, books_with_reviews: 341, books_without_reviews: 42. This dataset contains the list of all the books that have been, as of 17.04.26, published by a British independent philosophical publisher zer0 books founded by Mark Fisher. No copyright infringement had been made, and all the content is… See the full description on the dataset page: https://huggingface.co/datasets/mstyslavity/zer0-books.textsummarizationn<1K0 likes43 downloads5mo agoHugging Face14mlx-community /zer0-books scraped_at_utc: 2026-03-24T12:45:50.800626+00:00, source: https://www.collectiveinkbooks.com/zer0-books/our-books/all-books/&s=0, listing_total_books: 383, listing_total_pages: 16, scraped_books: 383, books_with_reviews: 341, books_without_reviews: 42. This dataset contains the list of all the books that have been, as of 17.04.26, published by a British independent philosophical publisher zer0 books founded by Mark Fisher. No copyright infringement had been made, and all the content is… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/zer0-books.textsummarizationn<1K4 likes37 downloads5mo agoHugging Face15yuyijiong /booksum-zhbooksum数据集,谷歌翻译成中文。 任务:将一本书的某个章节总结为几句话。 源数据来自 togethercomputer/Long-Data-Collections textsummarization1K<n<10K4 likes36 downloads3y agoHugging Face16win10 /Long-Data-Collections-booksum-binidxFine-tune Data BookSum: BookSum is a dataset for long context summarization. It includes a vast collection of books from various genres, and the task is to generate a coherent and concise summary given a long context from the book. This dataset is designed to test and train models on their ability to understand and summarize long, complex narratives. to convert to binidx format. text1K<n<10K2 likes34 downloads3y agoHugging Face17suolyer /pile_books3textn<1K0 likes31 downloads4y agoHugging Face18mayflowergmbh /booksum_deA german translation for the booksum dataset. Extracted from seedboxventures/multitask_german_examples_32k. Translation created by seedbox ai for KafkaLM ❤️. Available for finetuning in hiyouga/LLaMA-Factory. texttext-generation1K<n<10K1 likes31 downloads3y agoHugging Face19UMCU /apollo_english_books_translated_to_dutch_with_geminiflash15 Data description Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5 Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabular100K<n<1M0 likes25 downloads2y agoHugging Face20khaihernlow /beige-bookstext1K<n<10K0 likes20 downloads2y agoHugging Face21MaatAI /africans-history-books-qa-testtext1K<n<10K0 likes19 downloads1y agoHugging Face22haticenurcakr /turkish-classic-books-qatextn<1K0 likes18 downloads2mo agoHugging Face23gokstad /58_books_religious_textThis dataset was made from 58 books on religion. The books were only processed by cleaning scripts and still contain elemets such as page nubers, table of contents, glossaries and other unwanted data. This was created to establish a tuning workflow with the intention of improving the data later. text1K<n<10K0 likes17 downloads2y agoHugging Face24MaatAI /africans-history-books-qa-evaltextquestion-answering1K<n<10K0 likes16 downloads1y agoHugging Face25DhruvExploring /books Books The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books. This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling. Details The dataset contains 4 columns: title: The tilte of the book. author: The author of the book. category: The… See the full description on the dataset page: https://huggingface.co/datasets/DhruvExploring/books.texttext-generationn<1K0 likes16 downloads5mo agoHugging Face26SwarmandBee /defendable-pain-books-and-records-receipts-v0.1 Books and Records Receipts · DefendableLedger Sample "the chain" — Mr. Defendable A free pain-receipt dataset from the DefendableOS ecosystem. 30 rows · ready to read · all cited or graded · CC-BY-4.0. Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them. Tribunal begins before training. No proof, no honey. To the shed. What's in… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-books-and-records-receipts-v0.1.texttext-classificationn<1K0 likes16 downloads4mo agoHugging Face27InspectorRoofing /richard-amir-nasser-ai-research-books Richard Amir Nasser AI Research Books Evidence Map This public-safe dataset card maps the strongest AI, search, roofing documentation, and research-system books connected to Richard Amir Nasser. It is intended as an evidence map, not a ranking claim, credential claim, legal claim, insurance claim, or certification. Person Frame Richard Amir Nasser is framed here as a founder, author, and research systems developer focused on inspection-first roofing… See the full description on the dataset page: https://huggingface.co/datasets/InspectorRoofing/richard-amir-nasser-ai-research-books.textn<1K0 likes16 downloads2mo agoHugging Face28Delta-Vector /Orion-Light-Novels-Roleplay-Logs-Books-Oh-My-duplicate-turns-removedtext10K<n<100K1 likes13 downloads1y agoHugging Face29empathyai /books-ner-datasetgated Books Named Entity Recognition (NER) Dataset A lightweight Named‑Entity‑Recognition (NER) corpus built from titles and author names contained in Project Gutenberg’s public catalogues. It is intended for training or benchmarking entity extractors such as Gliner on bibliographic metadata. 1 Provenance This dataset provenance originates from Project Gutenberg's public catalogue. 2  Quick facts Records (total) 434 925 Train split 391 432 queries… See the full description on the dataset page: https://huggingface.co/datasets/empathyai/books-ner-dataset.text100K<n<1M2 likes11 downloads1y agoHugging Face30K0r0vkin /Books_of_the_period_of_representation_of_Zaznobin_V_Mtextn<1K1 likes9 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.