datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-books
Arabic Books
Dataset Summary
The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer.
This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.pre_1929_books_filtered
Pre-1929 Books
Description
Books published in the US before 1929 passed into the public domain on January 1, 2024.
We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929.
The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries.
The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.LoC-PD-Books
Library of Congress Public Domain Books (English)
This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic.
Curation method
The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books.
Dataset summary
The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.pre_1929_books
Pre-1929 Books
Description
Books published in the US before 1929 passed into the public domain on January 1, 2024.
We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929.
The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries.
The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books.Sci-Fi-Books-gutenberg
Gutenberg Sci-Fi Book Dataset
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
Data Format
The dataset is provided in CSV format. Each record represents a book and includes the following fields:
ID: A unique identifier for the book.
Title: The title of the book.
Author: The author(s) of the book.
Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.TamilNadu-State-board-books-2025-Tamil-and-english-versions
Tamil Nadu School Textbooks — Tamil and English Structured Text
This dataset contains text extracted from 312 Tamil Nadu State Board school
textbooks for Standards 1–12. It covers Tamil- and English-medium books and
provides each retained book in two forms:
structured JSON with book metadata, ordered sections, typed content blocks,
source references, extraction statistics, and curation provenance;
Markdown for reading, inspection, and downstream text processing.
The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.booksum-complete-cleaned
Description:
This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization
.
This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.Chinese-Psychology-Books
免责声明与使用须知 (Disclaimer and Usage Notice)
数据集内容
本数据集包含从互联网上多个来源收集的 中文心理学电子书 的集合。
许可证
本数据集的组织结构、汇编方式以及由维护者添加的任何元数据或注释根据 知识共享署名-非商业性使用 4.0 国际许可协议 (Creative Commons Attribution-NonCommercial 4.0 International License - CC BY-NC 4.0) 提供。这意味着您可以基于非商业目的分享和修改这部分内容,但必须给出适当的署名。
请注意:此 CC BY-NC 4.0 许可证不适用于数据集中包含的原始电子书文件本身。
版权声明
数据集中包含的个别电子书文件极有可能受到版权法保护,其版权归各自的作者、出版商或其他版权所有者所有。
数据集维护者不拥有这些电子书的版权。
这些电子书的来源多样且零散,部分来源可能难以追溯。
使用限制与责任… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Psychology-Books.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.the_pile_books3This dataset is Shawn Presser's work and is part of EleutherAi/The Pile dataset. This dataset contains all of bibliotik in plain .txt form, aka 197,000 books processed in exactly the same way as did for bookcorpusopen (a.k.a. books1). seems to be similar to OpenAI's mysterious "books2" dataset referenced in their papers. Unfortunately OpenAI will not give details, so we know very little about any differences. People suspect it's "all of libgen", but it's purely conjecture.uz-books
Dataset Card for BookCorpus
Dataset Summary
In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively.
Please refer to our blogpost and paper (Coming soon!) for further details.
To load and use dataset, run this script:… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-books.uz-books-v2
Dataset Card for UzBooks V2
Dataset Summary
UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits:
Split
Description
Examples
lat
Fully Latin-transliterated version
38,339
cyr
Fully Cyrillic-transliterated version
38,339
What's New in V2?
OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR
Cleaner Text: Google OCR produces far fewer recognition… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.portuguese-classic-books-adapted-to-modern-portuguese-brOkay, here is the improved and expanded text translated into American English, including the corrected citation format.
Classic Portuguese Language Books Adapted to Modern Brazilian Portuguese
Detailed Dataset Description
This dataset presents a unique collection of texts derived from classic books of Portuguese language literature, with a strong representation of Brazilian authors. All selected works are in the public domain and were originally sourced from… See the full description on the dataset page: https://huggingface.co/datasets/bobboyms/portuguese-classic-books-adapted-to-modern-portuguese-br.Hindawi-Books-dataset
Dataset Card for "Hindawi Books Dataset"
Hindawi Books Dataset is a large collection of more than 3000 books written in Modern Standard Arabic.
Dataset Description
Hindawi Books Dataset offers a rich and diverse collection of literary works, covering various topics and genres, all written in Modern Standard Arabic. The dataset includes information about each book, such as the title, author name, book abstract, and a link to access the complete text online. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/alielfilali01/Hindawi-Books-dataset.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.telugu_booksThis dataset is created by scraping telugu novels from teluguone.com this dataset can be used for nlp tasks like topic modeling, word embeddings, transfer learning etctask1650_opus_books_en-fi_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1650_opus_books_en-fi_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1650_opus_books_en-fi_translation.books2-1.2-lite
Dataset Card for books2-1.2-lite
Dataset Summary
books2-1.2 is a unofficial replication of rumors for CAI's books2 dataset.
Note
This dataset is meant for experimentation for others to try and provide feedback.As such this version is not compelete.
Contributions
KaraKaraWitch [Data gatering]
Anonymous [Compute provider]
books
Books
The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books.
This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling.
Details
The dataset contains 4 columns:
title: The tilte of the book.
author: The author of the book.
category: The… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/books.wiki-ru-en-news-booksptbr-books-publicos
PT-BR Public-Domain Books
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice.
Source and collection method
Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.Books-pdfArabic-books-and-research-dataset
Arabic reserach and books dataset (ARABD)
This dataset is an extracted cleaned text from more than 60K word files with unique arabic texts never published before.
Dataset diversity
the dataset is diverse from all kind of islamic research: [feqh, hadeeth, tafseer, tahqeeq, ... etc], from new written research to a manuscirpts.
dataset size
the dataset was more than 11GB but after cleaning (pre-processing) it becase a straight 10GB with less noisy chars.… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Arabic-books-and-research-dataset.Book_Summary_Chinese
中文图书总结数据集
每个样本包含:
图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。
rp_books-en
Dataset Card for "rp_books-en"
Filtering/cleaning on the 'red pajama books' subset of togethercomputer/Long-Data-Collections
The default config:
Dataset({
features: ['meta', 'text'],
num_rows: 26372
})
token count
default
GPT-4 tiktoken token count:
token_count
count 2.637200e+04
mean 1.009725e+05
std 1.161315e+05
min 3.811000e+03
25% 3.752750e+04
50% 7.757950e+04
75% 1.294130e+05
max 8.687685e+06
Total count: 2662.85 M… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/rp_books-en.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.rock-pollub-pl-books
ROCK Politechnika Lubelska PL books
A fail-closed Polish academic-book subset extracted from the official ROCK repository of Lublin University of Technology.
Retained books: 8
Text characters: 3,380,937
Tokens: 1,177,858 (cl100k_base proxy)
Author coverage: 100.0%
License: CC BY-SA 4.0, confirmed for every retained item and matched PDF bitstream
Source period: 2023-2026
The acquisition target was 20 books, but only eight passed the conservative per-file rights gate. The other… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/rock-pollub-pl-books.thai-tnhc2-books
Thai TNHC2 Books
This dataset collect all books from TNHC2 corpus.
We clean the dataset to use text to pretraining model and nlp task.
All books: 353 books
License: CC-0
TNHC2 Dataset (Original) have many a lots of details (chapter, author's detail and more). The dataset is clean to pretraining model and nlp task.
TNHC2 coepus is a Thai old books corpus that all books are copyright expired in Thai law (50 years after the author's death).
TNHC2 Dataset (Original):… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-tnhc2-books.
