CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MohamedRashad /arabic-books Arabic Books Dataset Summary The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.texttext-generation1K<n<10K3 likes34k downloads2y agoHugging Face02storytracer /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.tabulartext-generation100K<n<1M191 likes4.6k downloads3y agoHugging Face03PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes2.5k downloads3y agoHugging Face04common-pile /pre_1929_books_filtered Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.texttext-generation100K<n<1M2 likes1.8k downloads1y agoHugging Face05storytracer /LoC-PD-Books Library of Congress Public Domain Books (English) This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic. Curation method The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books. Dataset summary The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.tabulartext-generation10K<n<100K42 likes1.3k downloads3y agoHugging Face06common-pile /pre_1929_books Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books.texttext-generation100K<n<1M8 likes1.1k downloads1y agoHugging Face07stevez80 /Sci-Fi-Books-gutenberg Gutenberg Sci-Fi Book Dataset This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing. Data Format The dataset is provided in CSV format. Each record represents a book and includes the following fields: ID: A unique identifier for the book. Title: The title of the book. Author: The author(s) of the book. Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.texttext-generation1K<n<10K12 likes591 downloads3y agoHugging Face08Joemn /TamilNadu-State-board-books-2025-Tamil-and-english-versions Tamil Nadu School Textbooks — Tamil and English Structured Text This dataset contains text extracted from 312 Tamil Nadu State Board school textbooks for Standards 1–12. It covers Tamil- and English-medium books and provides each retained book in two forms: structured JSON with book metadata, ordered sections, typed content blocks, source references, extraction statistics, and curation provenance; Markdown for reading, inspection, and downstream text processing. The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.texttext-retrievaln<1K1 likes573 downloads2mo agoHugging Face09ubaada /booksum-complete-cleaned Description: This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization . This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.textsummarization1K<n<10K23 likes455 downloads2y agoHugging Face10Mxode /Chinese-Psychology-Books 免责声明与使用须知 (Disclaimer and Usage Notice) 数据集内容 本数据集包含从互联网上多个来源收集的 中文心理学电子书 的集合。 许可证 本数据集的组织结构、汇编方式以及由维护者添加的任何元数据或注释根据 知识共享署名-非商业性使用 4.0 国际许可协议 (Creative Commons Attribution-NonCommercial 4.0 International License - CC BY-NC 4.0) 提供。这意味着您可以基于非商业目的分享和修改这部分内容,但必须给出适当的署名。 请注意:此 CC BY-NC 4.0 许可证不适用于数据集中包含的原始电子书文件本身。 版权声明 数据集中包含的个别电子书文件极有可能受到版权法保护,其版权归各自的作者、出版商或其他版权所有者所有。 数据集维护者不拥有这些电子书的版权。 这些电子书的来源多样且零散,部分来源可能难以追溯。 使用限制与责任… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Psychology-Books.texttext-generationn<1K9 likes437 downloads1y agoHugging Face11BrightData /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.tabulartext-classification1M<n<10M20 likes420 downloads2y agoHugging Face12defunct-datasets /the_pile_books3This dataset is Shawn Presser's work and is part of EleutherAi/The Pile dataset. This dataset contains all of bibliotik in plain .txt form, aka 197,000 books processed in exactly the same way as did for bookcorpusopen (a.k.a. books1). seems to be similar to OpenAI's mysterious "books2" dataset referenced in their papers. Unfortunately OpenAI will not give details, so we know very little about any differences. People suspect it's "all of libgen", but it's purely conjecture.text-generation100K<n<1M153 likes342 downloads3y agoHugging Face13murodbek /uz-books Dataset Card for BookCorpus Dataset Summary In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively. Please refer to our blogpost and paper (Coming soon!) for further details. To load and use dataset, run this script:… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-books.texttext-generation10K<n<100K23 likes315 downloads1y agoHugging Face14tahrirchi /uz-books-v2 Dataset Card for UzBooks V2 Dataset Summary UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits: Split Description Examples lat Fully Latin-transliterated version 38,339 cyr Fully Cyrillic-transliterated version 38,339 What's New in V2? OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR Cleaner Text: Google OCR produces far fewer recognition… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.texttext-generation10K<n<100K5 likes247 downloads6mo agoHugging Face15bobboyms /portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format. Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese Detailed Dataset Description This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.texttext-generation10K<n<100K0 likes170 downloads1y agoHugging Face16alielfilali01 /Hindawi-Books-dataset Dataset Card for "Hindawi Books Dataset" Hindawi Books Dataset is a large collection of more than 3000 books written in Modern Standard Arabic. Dataset Description Hindawi Books Dataset offers a rich and diverse collection of literary works, covering various topics and genres, all written in Modern Standard Arabic. The dataset includes information about each book, such as the title, author name, book abstract, and a link to access the complete text online. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/alielfilali01/Hindawi-Books-dataset.texttext-generation10K<n<100K14 likes152 downloads3y agoHugging Face17PleIAs /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K4 likes151 downloads3y agoHugging Face18community-datasets /telugu_booksThis dataset is created by scraping telugu novels from teluguone.com this dataset can be used for nlp tasks like topic modeling, word embeddings, transfer learning etctext-generationn<1K4 likes144 downloads3y agoHugging Face19Lots-of-LoRAs /task1650_opus_books_en-fi_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1650_opus_books_en-fi_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1650_opus_books_en-fi_translation.texttext-generation1K<n<10K0 likes118 downloads2y agoHugging Face20RyokoExtra /books2-1.2-lite Dataset Card for books2-1.2-lite Dataset Summary books2-1.2 is a unofficial replication of rumors for CAI's books2 dataset. Note This dataset is meant for experimentation for others to try and provide feedback.As such this version is not compelete. Contributions KaraKaraWitch [Data gatering] Anonymous [Compute provider] text-generation100K<n<1M8 likes113 downloads3y agoHugging Face21IsmaelMousa /books Books The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books. This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling. Details The dataset contains 4 columns: title: The tilte of the book. author: The author of the book. category: The… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/books.texttext-generationn<1K5 likes99 downloads2y agoHugging Face22ai-allforever /wiki-ru-en-news-bookstexttext-generation1M<n<10M2 likes99 downloads10mo agoHugging Face23dataseek /ptbr-books-publicos PT-BR Public-Domain Books Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice. Source and collection method Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.tabulartext-generation10K<n<100K0 likes92 downloads5mo agoHugging Face24Diablo99t /Books-pdfdocumenttext-generationn<1K0 likes92 downloads6d agoHugging Face25riotu-lab /Arabic-books-and-research-dataset Arabic reserach and books dataset (ARABD) This dataset is an extracted cleaned text from more than 60K word files with unique arabic texts never published before. Dataset diversity the dataset is diverse from all kind of islamic research: [feqh, hadeeth, tafseer, tahqeeq, ... etc], from new written research to a manuscirpts. dataset size the dataset was more than 11GB but after cleaning (pre-processing) it becase a straight 10GB with less noisy chars.… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Arabic-books-and-research-dataset.texttext-generation10K<n<100K6 likes88 downloads2y agoHugging Face26yuyijiong /Book_Summary_Chinese 中文图书总结数据集 每个样本包含: 图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。 texttext-generationn<1K27 likes86 downloads3y agoHugging Face27BEE-spoke-data /rp_books-en Dataset Card for "rp_books-en" Filtering/cleaning on the 'red pajama books' subset of togethercomputer/Long-Data-Collections The default config: Dataset({ features: ['meta', 'text'], num_rows: 26372 }) token count default GPT-4 tiktoken token count: token_count count 2.637200e+04 mean 1.009725e+05 std 1.161315e+05 min 3.811000e+03 25% 3.752750e+04 50% 7.757950e+04 75% 1.294130e+05 max 8.687685e+06 Total count: 2662.85 M… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/rp_books-en.texttext-generation100K<n<1M1 likes83 downloads9mo agoHugging Face28codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes79 downloads11mo agoHugging Face29PiotrSty /rock-pollub-pl-books ROCK Politechnika Lubelska PL books A fail-closed Polish academic-book subset extracted from the official ROCK repository of Lublin University of Technology. Retained books: 8 Text characters: 3,380,937 Tokens: 1,177,858 (cl100k_base proxy) Author coverage: 100.0% License: CC BY-SA 4.0, confirmed for every retained item and matched PDF bitstream Source period: 2023-2026 The acquisition target was 20 books, but only eight passed the conservative per-file rights gate. The other… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/rock-pollub-pl-books.texttext-generationn<1K0 likes78 downloads16d agoHugging Face30pythainlp /thai-tnhc2-books Thai TNHC2 Books This dataset collect all books from TNHC2 corpus. We clean the dataset to use text to pretraining model and nlp task. All books: 353 books License: CC-0 TNHC2 Dataset (Original) have many a lots of details (chapter, author's detail and more). The dataset is clean to pretraining model and nlp task. TNHC2 coepus is a Thai old books corpus that all books are copyright expired in Thai law (50 years after the author's death). TNHC2 Dataset (Original):… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-tnhc2-books.texttext-generationn<1K0 likes73 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.