datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining_v1-omega_booksUS-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.booksum
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization
Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev
Introduction
The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases.
While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.LoC-PD-Books
Library of Congress Public Domain Books (English)
This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic.
Curation method
The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books.
Dataset summary
The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.goodreads-books
Goodreads Books Metadata
Dataset Description
Goodreads Books Metadata is a structured dataset of book records scraped directly from Goodreads, a social platform for book readers and recommendations.
The dataset is collected via an ongoing, resumable crawl and contains rich metadata per book: bibliographic information, crowd-sourced ratings, contributor (author/illustrator/editor/etc.) details enriched with author-level popularity stats, genre tags, series… See the full description on the dataset page: https://huggingface.co/datasets/pfaha/goodreads-books.IA-booksAmazon-Reviews-2023-Books-Review
Amazon Reviews 2023 (Books Only)
This is a subset of Amazon Review 2023 dataset. Please visit amazon-reviews-2023.github.io/ for more details, loading scripts, and preprocessed benchmark files.
[April 18, 2024] Update
This dataset was created and pushed for the first time.
This is a large-scale Amazon Reviews dataset, collected in 2023 by McAuley Lab, and it includes rich features such as:
User Reviews (ratings, text, helpfulness votes, etc.);
Item Metadata (descriptions… See the full description on the dataset page: https://huggingface.co/datasets/cogsci13/Amazon-Reviews-2023-Books-Review.Amazon-Reviews-2023-Books-Meta
Amazon Reviews 2023 (Books Only)
This is a subset of Amazon Review 2023 dataset. Please visit amazon-reviews-2023.github.io/ for more details, loading scripts, and preprocessed benchmark files.
[April 18, 2024] Update
This dataset was created and pushed for the first time.
This is a large-scale Amazon Reviews dataset, collected in 2023 by McAuley Lab, and it includes rich features such as:
User Reviews (ratings, text, helpfulness votes, etc.);
Item Metadata (descriptions… See the full description on the dataset page: https://huggingface.co/datasets/cogsci13/Amazon-Reviews-2023-Books-Meta.AIRBOT_MMK2_organize_and_place_books
AIRBOT_MMK2_organize_and_place_books
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_organize_and_place_books.AIRBOT_MMK2_organize_books
AIRBOT_MMK2_organize_books
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_organize_books.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.AIRBOT_MMK2_place_the_books
AIRBOT_MMK2_place_the_books
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
push
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_books.zlib-books-1k-50kinstitutional-books-hl-metadata
📚 Institutional Books: Harvard Library (metadata-only version)
Institutional Books is a growing corpus of public domain books. This release is comprised of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative. IDI Terms of Use for Early-Access.
983K books, published largely in the 19th and 20th centuries
242B o200k_base tokens
386M pages of text, available in both original… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-metadata.google-books-ngram-pos
Google Books Ngram — POS-tagged & cleaned
A cleaned, analysis-ready slice of the Google Books Ngram corpus v3
(20200217, English) with part-of-speech tags preserved, packaged as
Parquet for easy use with 🤗 datasets, pandas, DuckDB, or Polars.
This powers the POS / regex Ngram Viewer — an ngram viewer that supports
part-of-speech template queries (love *_NOUN) and regex (/ousness$/_NOUN),
patterns the official Google viewer cannot express.
Configs
config… See the full description on the dataset page: https://huggingface.co/datasets/dhruv-anand-aintech/google-books-ngram-pos.zlib-books-1k-500kUkrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.project-madurai-booksProject Madurai Books Text Dataset
This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website.
Dataset Details
You can see a table above called "Meta Data", which is just an info table.
You can't able to preview the "Source Data" table, due to it being about 300MB.
[Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.books3_lowgrade_paragraphs
Dataset Card for "books3_lowgrade_paragraphs"
the_pile books3, books with smog grade difficulty estimate between 6.6 or and 7.1. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc.
For easier books, see books3_basic_paragraphs
media-metadata-openlibrary-books
TigreGotico/media-metadata-openlibrary-books
Rich entity dataset scraped by metadatarr
scraper openlibrary_books.
Rows: 4,098,190
Fields
olid
title
subtitle
authors
author_key
first_publish_year
subjects
isbn_10
isbn_13
publisher
language
number_of_pages_median
ebook_access
has_fulltext
edition_count
cover_i
Source
Generated by scrapers/openlibrary_books.py. See the metadatarr repo for the full
pipeline and scraper source code.
books-ocrThe data is composed of 5 json files. Each file is an array of these objects:
{
"pageId": "clzo00pu002a1cmtc4dws1nzr",
"url": "https://assets.usul.ai/ocr-pages/clzo00pu002a1cmtc4dws1nzr.png",
"pdfPageNumber": 2344,
"ocrOutput": "٢٣٤٣\nعيون المجالس\n١٠٢٧ - مسألة: إذا اشترى المشتري عبدًا أو أمة أو سلعة من السلع\nفحدث عنده عیب ثم وجد به عيبًا عند البائع . ١٤٦٧\n١٠٢٨ - مسألة: إذا ابتاع الرجل شيئًا فوجد به عيبًا فقال: فسخت ١٤٦٨\nالبيع.\n١٠٢٩ - مسألة: عندنا أن العبد ملك لا يساوي الحر فيه.… See the full description on the dataset page: https://huggingface.co/datasets/seemorg/books-ocr.premodern-japanese-books-lm-corpus
Premodern Japanese Books LM Corpus
日本語
概要
日本古典籍統一データセットの言語モデル学習用本文ビュー v0.2.0 です。lm-curated v0.2.0から、本文採用対象とした1,598文書を収録しています。文書の本文はcontent列に入り、文書単位の論理分割はsplit列に記録しています。
収録範囲
kouigenji
ndl-minhon-ocrdataset
yatanavi
利用上の注意
Hugging Face上の物理splitはtrain一つです。split列にtrain、validation、testの論理分割を保持しています。
NDL Minhonでは角括弧の記号だけを削除し、角括弧内部の文字は保持しています。
やたナビでは読み仮名、異読、校訂注、その他の補助表記を除去しています。
利用条件と帰属表示はNOTICE.mdを確認してください。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-lm-corpus.ptbr-books-publicos
PT-BR Public-Domain Books
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice.
Source and collection method
Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tamil_books_na
Tamil Books Dataset
A collection of Tamil books in digital format for natural language processing and research purposes.
Dataset Description
This dataset contains Tamil books converted from various sources into a structured format suitable for NLP tasks.
Features
id: Unique identifier
title: Title of the book
author: Author of the book
content: The main text content of the book
Source Acknowledgement
I get Tamil books from this site:… See the full description on the dataset page: https://huggingface.co/datasets/Naveen934/tamil_books_na.whole-books
Whole Books
Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book.
Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large
majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below;
nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are
documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.books3_basic_paragraphs
Dataset Card for "books3_basic_paragraphs"
the_pile books3, books with smog grade difficulty estimate of 6.5 or under. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc.
premodern-japanese-books-source-crosswalk
Premodern Japanese Books Source Crosswalk
日本語
概要
日本古典籍統一データセット v0.1.0 の公開用Viewです。固定済み内部Releaseから、再配布と機械学習利用が許可された行だけを収録しています。収録行数は 17,570,138 行、収録SourceDataset数は 7 件です。
用途
日本語歴史資料の研究、検索、OCRまたは言語モデル用データ処理に利用できます。個々の行には採用したCanonical ID、権利判定、必要な帰属を保持しています。
権利
単一のライセンス値は全行の条件を表しません。必ず NOTICE.md と行単位の権利列を確認してください。
限界
v0.1.0 は監査対象52候補のうち、固定入力が成立した21 SourceDatasetを対象とする段階公開です。内容の正確性、外部参照の永続性、特定用途への適合性を保証しません。… See the full description on the dataset page: https://huggingface.co/datasets/Kotomiya07/premodern-japanese-books-source-crosswalk.Shamela_Books_info
Shamela Books information
This dataset contains structured metadata for 8,492 books sourced from the Shamela Library, with enhancements for clarity, consistency, and usability. It is intended to support NLP, bibliographic research, and digital humanities efforts involving Arabic texts.For full books text dataset please check shamela_books_text
Dataset Features
The dataset includes the following cleaned and standardized features:
Unification of Author Names: Author… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/Shamela_Books_info.
