CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes381k downloads2y agoHugging Face02MohamedRashad /arabic-books Arabic Books Dataset Summary The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.texttext-generation1K<n<10K3 likes34k downloads2y agoHugging Face03chcaa /kb-books open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish Library published between 1750 and 1930. The dataset has each page of each document in image and text format. The text was extracted with OCR. The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.image1M<n<10M4 likes20k downloads10mo agoHugging Face04Helsinki-NLP /opus_books Dataset Card for OPUS Books Dataset Summary This is a collection of copyright free books aligned by Andras Farkas, which are available from http://www.farkastranslations.com/bilingual_books.php Note that the texts are rather dated due to copyright issues and that some of them are manually reviewed (check the meta-data at the top of the corpus files in XML). The source is multilingually aligned, which is available from http://www.farkastranslations.com/bilingual_books.php.… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/opus_books.texttranslation1M<n<10M96 likes10k downloads2y agoHugging Face05institutional /institutional-books-hl-visual-elementsgated 📚 Institutional Books: Harvard Library — Visual Elements 22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset. 22,622,060 visual elements extracted from 983,004 volumes 766,992,447 o200k_base tokens in AI-generated captions 6 high-level classes of visual elements organized in splits 5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.image10M<n<100M1 likes8.2k downloads1mo agoHugging Face06tiendung /vi-books_extractionll books_text_final/ | wc -l => 13968 text After filter (not vi, too small) wc -l vi-books.jsonl => 12832 text => !!! Cần lọc truyện từ dtv ra !!! xzcat dtv_ebooks_details.jsonl.xz | wc -l # => 13486 ll books_text_final/ | grep dtv_ebooks_ | wc -l # => 3535 từ dtv lấy các whitelist categories tại dtv_categories_whitelist.jsonl lần theo từng ebook một trong dtv_ebooks_details.jsonl.xz trường {"cat": "Self Help - Khởi nghiệp", "cat_url":… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/vi-books_extraction.text0 likes5.5k downloads3y agoHugging Face07storytracer /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.tabulartext-generation100K<n<1M191 likes4.6k downloads3y agoHugging Face08kmfoda /booksum BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev Introduction The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases. While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.tabular10K<n<100K80 likes3.9k downloads4y agoHugging Face09muse-bench /MUSE-Books MUSE-Books MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-Books.textn<1K3 likes3.8k downloads2y agoHugging Face10tiendung /vi-books_dtv-ebook.com git lfs track "*.epub" git lfs track "*.mobi" git lfs track "*.prc" git lfs track "*.txt" git lfs track "*.json" git lfs track "*.jsonl" git lfs track "*.rar" git lfs track "*.zip" git lfs track "*.doc" git lfs track "*.docx" git lfs track "*.pdf" git lfs track "*.azw3" du -sh download/* > _size_downloads.txt documentn<1K0 likes3.5k downloads3y agoHugging Face11PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes2.5k downloads3y agoHugging Face12cloudaocr /arabic-synthetic-scanned-booksdocument1K<n<10K1 likes2.4k downloads23d agoHugging Face13freococo /ocr_arabic_books Arabic OCR Books Dataset (ocr_arabic_books) This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions. 📚 Master Book Inventory Total running pages in repository: 181,427 # Book Name (English) Book Name (Arabic) Subset / Config Name Page Count Image Index Range 1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.imageimage-to-text100K<n<1M4 likes2.2k downloads3mo agoHugging Face14PCNTechnologyEnterprise /Booksdocument1M<n<10M0 likes2k downloads2mo agoHugging Face15ada-datadruids /booksummaries_cleanedtext10K<n<100K0 likes2k downloads2y agoHugging Face16freococo /synth_shamela_ocr_arabic_books Synthetic Arabic Books Dataset Structured book pages rendered dynamically with style, font, and degradation variations. imageimage-to-text1M<n<10M2 likes1.9k downloads2mo agoHugging Face17common-pile /pre_1929_books_filtered Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.texttext-generation100K<n<1M2 likes1.8k downloads1y agoHugging Face18emozilla /Long-Data-Collections-Pretrain-Without-Books Dataset Card for "Long-Data-Collections-Pretrain-Without-Books" Paraquet version of the pretrain split of togethercomputer/Long-Data-Collections WITHOUT books Statistics (in # of characters): total_len: 236088622215, average_len: 25159.041601590307 text1M<n<10M4 likes1.5k downloads3y agoHugging Face19davanstrien /bl_books_flickr_fullimage1M<n<10M1 likes1.5k downloads4y agoHugging Face20storytracer /LoC-PD-Books Library of Congress Public Domain Books (English) This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic. Curation method The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books. Dataset summary The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.tabulartext-generation10K<n<100K42 likes1.3k downloads3y agoHugging Face21pfaha /goodreads-books Goodreads Books Metadata Dataset Description Goodreads Books Metadata is a structured dataset of book records scraped directly from Goodreads, a social platform for book readers and recommendations. The dataset is collected via an ongoing, resumable crawl and contains rich metadata per book: bibliographic information, crowd-sourced ratings, contributor (author/illustrator/editor/etc.) details enriched with author-level popularity stats, genre tags, series… See the full description on the dataset page: https://huggingface.co/datasets/pfaha/goodreads-books.tabulartabular-regression100K<n<1M0 likes1.1k downloads9h agoHugging Face22common-pile /pre_1929_books Pre-1929 Books Description Books published in the US before 1929 passed into the public domain on January 1, 2024. We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929. The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries. The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books.texttext-generation100K<n<1M8 likes1.1k downloads1y agoHugging Face23izumi-lab /open-text-books Dataset Card for "open-text-books" More Information needed text100K<n<1M18 likes903 downloads3y agoHugging Face24SaylorTwift /the_pile_books3_minus_gutenberg Dataset Card for "the_pile_books3_minus_gutenberg" More Information needed text100K<n<1M15 likes857 downloads4y agoHugging Face25marianna13 /IA-bookstabular1M<n<10M0 likes619 downloads4y agoHugging Face26stevez80 /Sci-Fi-Books-gutenberg Gutenberg Sci-Fi Book Dataset This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing. Data Format The dataset is provided in CSV format. Each record represents a book and includes the following fields: ID: A unique identifier for the book. Title: The title of the book. Author: The author(s) of the book. Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.texttext-generation1K<n<10K12 likes591 downloads3y agoHugging Face27cogsci13 /Amazon-Reviews-2023-Books-Meta Amazon Reviews 2023 (Books Only) This is a subset of Amazon Review 2023 dataset. Please visit amazon-reviews-2023.github.io/ for more details, loading scripts, and preprocessed benchmark files. [April 18, 2024] Update This dataset was created and pushed for the first time. This is a large-scale Amazon Reviews dataset, collected in 2023 by McAuley Lab, and it includes rich features such as: User Reviews (ratings, text, helpfulness votes, etc.); Item Metadata (descriptions… See the full description on the dataset page: https://huggingface.co/datasets/cogsci13/Amazon-Reviews-2023-Books-Meta.tabular1M<n<10M9 likes582 downloads2y agoHugging Face28Joemn /TamilNadu-State-board-books-2025-Tamil-and-english-versions Tamil Nadu School Textbooks — Tamil and English Structured Text This dataset contains text extracted from 312 Tamil Nadu State Board school textbooks for Standards 1–12. It covers Tamil- and English-medium books and provides each retained book in two forms: structured JSON with book metadata, ordered sections, typed content blocks, source references, extraction statistics, and curation provenance; Markdown for reading, inspection, and downstream text processing. The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.texttext-retrievaln<1K1 likes573 downloads2mo agoHugging Face29cogsci13 /Amazon-Reviews-2023-Books-Review Amazon Reviews 2023 (Books Only) This is a subset of Amazon Review 2023 dataset. Please visit amazon-reviews-2023.github.io/ for more details, loading scripts, and preprocessed benchmark files. [April 18, 2024] Update This dataset was created and pushed for the first time. This is a large-scale Amazon Reviews dataset, collected in 2023 by McAuley Lab, and it includes rich features such as: User Reviews (ratings, text, helpfulness votes, etc.); Item Metadata (descriptions… See the full description on the dataset page: https://huggingface.co/datasets/cogsci13/Amazon-Reviews-2023-Books-Review.tabular10M<n<100M1 likes572 downloads2y agoHugging Face30PleIAs /BDH-Books 🇪🇸 Biblioteca Digitale Hispanica - Books 🇪🇸 Biblioteca Digitale Hispanica-Books or BDH-Books is a large collection aiming to aggregate all Spanish books in the public domain coming from the Biblioteca Digitale Hispanica. Dataset summary The collection contains 139,932 individual titles mostly published in the 19th century and the first half of the 20th century, making up nearly 11 billion words (10,753,912,288 space-separated words). Curation method The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/BDH-Books.text100K<n<1M0 likes545 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.