datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.LoC-PD-Books
Library of Congress Public Domain Books (English)
This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic.
Curation method
The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books.
Dataset summary
The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.ptbr-books-publicos
PT-BR Public-Domain Books
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice.
Source and collection method
Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.whole-books
Whole Books
Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book.
Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large
majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below;
nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are
documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Chima207/Goodreads-Books.tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.LoC-PD-Books-preprocessed
LoC-PD-Books: preprocessed
This is the storytracer/LoC-PD-Books dataset with the following preprocessing steps:
apply clean-text package keeping casing and newlines
drop OCR garbled text in first few lines of each example
fix (most) 'hard' newlines w/ regex similar to gutenberg clean
'grade' first 512 tokens of each book with this quantized model; keep examples from labels clean (all) and mild gibberish w/ score 0.9 or higher
goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/bstarrs/goodreads-books.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/actixon/US-PD-Books.Books-General-Linux
Linux Books Dataset
Dataset Description
The Linux Books Dataset is a curated text dataset derived from Linux-related books and learning materials. It focuses on Linux system administration, cybersecurity, networking, shell scripting, and operating system fundamentals.The dataset is designed to support training and evaluation of NLP models for technical domains, especially cybersecurity-aware language models and Linux-focused assistants.
This dataset is suitable for both… See the full description on the dataset page: https://huggingface.co/datasets/DexopT/Books-General-Linux.
