CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes409k downloads2y agoHugging Face02storytracer /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.tabulartext-generation100K<n<1M191 likes4.5k downloads3y agoHugging Face03kmfoda /booksum BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev Introduction The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases. While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.tabular10K<n<100K80 likes3.9k downloads4y agoHugging Face04PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes2.5k downloads3y agoHugging Face05storytracer /LoC-PD-Books Library of Congress Public Domain Books (English) This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic. Curation method The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books. Dataset summary The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.tabulartext-generation10K<n<100K42 likes1.4k downloads3y agoHugging Face06pfaha /goodreads-books Goodreads Books Metadata Dataset Description Goodreads Books Metadata is a structured dataset of book records scraped directly from Goodreads, a social platform for book readers and recommendations. The dataset is collected via an ongoing, resumable crawl and contains rich metadata per book: bibliographic information, crowd-sourced ratings, contributor (author/illustrator/editor/etc.) details enriched with author-level popularity stats, genre tags, series… See the full description on the dataset page: https://huggingface.co/datasets/pfaha/goodreads-books.tabulartabular-regression100K<n<1M0 likes1.2k downloads9h agoHugging Face07marianna13 /IA-bookstabular1M<n<10M0 likes644 downloads4y agoHugging Face08cogsci13 /Amazon-Reviews-2023-Books-Review Amazon Reviews 2023 (Books Only) This is a subset of Amazon Review 2023 dataset. Please visit amazon-reviews-2023.github.io/ for more details, loading scripts, and preprocessed benchmark files. [April 18, 2024] Update This dataset was created and pushed for the first time. This is a large-scale Amazon Reviews dataset, collected in 2023 by McAuley Lab, and it includes rich features such as: User Reviews (ratings, text, helpfulness votes, etc.); Item Metadata (descriptions… See the full description on the dataset page: https://huggingface.co/datasets/cogsci13/Amazon-Reviews-2023-Books-Review.tabular10M<n<100M1 likes603 downloads2y agoHugging Face09cogsci13 /Amazon-Reviews-2023-Books-Meta Amazon Reviews 2023 (Books Only) This is a subset of Amazon Review 2023 dataset. Please visit amazon-reviews-2023.github.io/ for more details, loading scripts, and preprocessed benchmark files. [April 18, 2024] Update This dataset was created and pushed for the first time. This is a large-scale Amazon Reviews dataset, collected in 2023 by McAuley Lab, and it includes rich features such as: User Reviews (ratings, text, helpfulness votes, etc.); Item Metadata (descriptions… See the full description on the dataset page: https://huggingface.co/datasets/cogsci13/Amazon-Reviews-2023-Books-Meta.tabular1M<n<10M9 likes598 downloads2y agoHugging Face10RoboCOIN /AIRBOT_MMK2_organize_and_place_booksgated AIRBOT_MMK2_organize_and_place_books 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_organize_and_place_books.tabularrobotics1K<n<10K0 likes457 downloads9mo agoHugging Face11RoboCOIN /AIRBOT_MMK2_organize_booksgated AIRBOT_MMK2_organize_books 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_organize_books.tabularrobotics1K<n<10K0 likes452 downloads9mo agoHugging Face12BrightData /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.tabulartext-classification1M<n<10M20 likes423 downloads2y agoHugging Face13RoboCOIN /AIRBOT_MMK2_place_the_booksgated AIRBOT_MMK2_place_the_books 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: discover_robotics_aitbot_mmk2 | Codebase Version: v2.1 End-Effector Type: five_finger_hand 🏠 Scene Types This dataset covers the following scene types: home 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place push 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_books.tabularrobotics10K<n<100K0 likes375 downloads9mo agoHugging Face14marianna13 /zlib-books-1k-50ktabular100K<n<1M1 likes351 downloads3y agoHugging Face15institutional /institutional-books-hl-metadata 📚 Institutional Books: Harvard Library (metadata-only version) Institutional Books is a growing corpus of public domain books. This release is comprised of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative. IDI Terms of Use for Early-Access. 983K books, published largely in the 19th and 20th centuries 242B o200k_base tokens 386M pages of text, available in both original… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-metadata.tabular100K<n<1M17 likes264 downloads1mo agoHugging Face16dhruv-anand-aintech /google-books-ngram-pos Google Books Ngram — POS-tagged & cleaned A cleaned, analysis-ready slice of the Google Books Ngram corpus v3 (20200217, English) with part-of-speech tags preserved, packaged as Parquet for easy use with 🤗 datasets, pandas, DuckDB, or Polars. This powers the POS / regex Ngram Viewer — an ngram viewer that supports part-of-speech template queries (love *_NOUN) and regex (/ousness$/_NOUN), patterns the official Google viewer cannot express. Configs config… See the full description on the dataset page: https://huggingface.co/datasets/dhruv-anand-aintech/google-books-ngram-pos.tabulartext-classification1M<n<10M0 likes185 downloads3mo agoHugging Face17marianna13 /zlib-books-1k-500ktabular100K<n<1M0 likes183 downloads3y agoHugging Face18PleIAs /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K4 likes165 downloads3y agoHugging Face19mastergokul /project-madurai-booksProject Madurai Books Text Dataset This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website. Dataset Details You can see a table above called "Meta Data", which is just an info table. You can't able to preview the "Source Data" table, due to it being about 300MB. [Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.tabulartext-classification1K<n<10K0 likes129 downloads2y agoHugging Face20skeskinen /books3_lowgrade_paragraphs Dataset Card for "books3_lowgrade_paragraphs" the_pile books3, books with smog grade difficulty estimate between 6.6 or and 7.1. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc. For easier books, see books3_basic_paragraphs tabular10M<n<100M0 likes118 downloads3y agoHugging Face21LeData /media-metadata-openlibrary-books TigreGotico/media-metadata-openlibrary-books Rich entity dataset scraped by metadatarr scraper openlibrary_books. Rows: 4,098,190 Fields olid title subtitle authors author_key first_publish_year subjects isbn_10 isbn_13 publisher language number_of_pages_median ebook_access has_fulltext edition_count cover_i Source Generated by scrapers/openlibrary_books.py. See the metadatarr repo for the full pipeline and scraper source code. tabular1M<n<10M0 likes115 downloads3mo agoHugging Face22seemorg /books-ocrThe data is composed of 5 json files. Each file is an array of these objects: { "pageId": "clzo00pu002a1cmtc4dws1nzr", "url": "https://assets.usul.ai/ocr-pages/clzo00pu002a1cmtc4dws1nzr.png", "pdfPageNumber": 2344, "ocrOutput": "٢٣٤٣\nعيون المجالس\n١٠٢٧ - مسألة: إذا اشترى المشتري عبدًا أو أمة أو سلعة من السلع\nفحدث عنده عیب ثم وجد به عيبًا عند البائع . ١٤٦٧\n١٠٢٨ - مسألة: إذا ابتاع الرجل شيئًا فوجد به عيبًا فقال: فسخت ١٤٦٨\nالبيع.\n١٠٢٩ - مسألة: عندنا أن العبد ملك لا يساوي الحر فيه.… See the full description on the dataset page: https://huggingface.co/datasets/seemorg/books-ocr.image10K<n<100K6 likes104 downloads1y agoHugging Face23Kotomiya07 /premodern-japanese-books-lm-corpus Premodern Japanese Books LM Corpus 日本語 概要 日本古典籍統一データセットの言語モデル学習用本文ビュー v0.2.0 です。lm-curated v0.2.0から、本文採用対象とした1,598文書を収録しています。文書の本文はcontent列に入り、文書単位の論理分割はsplit列に記録しています。 収録範囲 kouigenji ndl-minhon-ocrdataset yatanavi 利用上の注意 Hugging Face上の物理splitはtrain一つです。split列にtrain、validation、testの論理分割を保持しています。 NDL Minhonでは角括弧の記号だけを削除し、角括弧内部の文字は保持しています。 やたナビでは読み仮名、異読、校訂注、その他の補助表記を除去しています。 利用条件と帰属表示はNOTICE.mdを確認してください。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-lm-corpus.tabular1K<n<10K0 likes98 downloads19d agoHugging Face24dataseek /ptbr-books-publicos PT-BR Public-Domain Books Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice. Source and collection method Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.tabulartext-generation10K<n<100K0 likes92 downloads5mo agoHugging Face25codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes78 downloads11mo agoHugging Face26Naveen934 /tamil_books_na Tamil Books Dataset A collection of Tamil books in digital format for natural language processing and research purposes. Dataset Description This dataset contains Tamil books converted from various sources into a structured format suitable for NLP tasks. Features id: Unique identifier title: Title of the book author: Author of the book content: The main text content of the book Source Acknowledgement I get Tamil books from this site:… See the full description on the dataset page: https://huggingface.co/datasets/Naveen934/tamil_books_na.tabular10K<n<100K1 likes68 downloads11mo agoHugging Face27ysngkil /whole-books Whole Books Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book. Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below; nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.tabulartext-generation100K<n<1M0 likes65 downloads21d agoHugging Face28skeskinen /books3_basic_paragraphs Dataset Card for "books3_basic_paragraphs" the_pile books3, books with smog grade difficulty estimate of 6.5 or under. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc. tabular1M<n<10M0 likes61 downloads3y agoHugging Face29Kotomiya07 /premodern-japanese-books-source-crosswalk Premodern Japanese Books Source Crosswalk 日本語 概要 日本古典籍統一データセット v0.1.0 の公開用Viewです。固定済み内部Releaseから、再配布と機械学習利用が許可された行だけを収録しています。収録行数は 17,570,138 行、収録SourceDataset数は 7 件です。 用途 日本語歴史資料の研究、検索、OCRまたは言語モデル用データ処理に利用できます。個々の行には採用したCanonical ID、権利判定、必要な帰属を保持しています。 権利 単一のライセンス値は全行の条件を表しません。必ず NOTICE.md と行単位の権利列を確認してください。 限界 v0.1.0 は監査対象52候補のうち、固定入力が成立した21 SourceDatasetを対象とする段階公開です。内容の正確性、外部参照の永続性、特定用途への適合性を保証しません。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-source-crosswalk.tabular10M<n<100M0 likes60 downloads20d agoHugging Face30MoMonir /Shamela_Books_info Shamela Books information This dataset contains structured metadata for 8,492 books sourced from the Shamela Library, with enhancements for clarity, consistency, and usability. It is intended to support NLP, bibliographic research, and digital humanities efforts involving Arabic texts.For full books text dataset please check shamela_books_text Dataset Features The dataset includes the following cleaned and standardized features: Unification of Author Names: Author… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/Shamela_Books_info.tabular1K<n<10K1 likes59 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.