datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.LoC-PD-Books
Library of Congress Public Domain Books (English)
This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic.
Curation method
The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books.
Dataset summary
The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.uz-books
Dataset Card for BookCorpus
Dataset Summary
In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively.
Please refer to our blogpost and paper (Coming soon!) for further details.
To load and use dataset, run this script:… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-books.uz-books-v2
Dataset Card for UzBooks V2
Dataset Summary
UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits:
Split
Description
Examples
lat
Fully Latin-transliterated version
38,339
cyr
Fully Cyrillic-transliterated version
38,339
What's New in V2?
OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR
Cleaner Text: Google OCR produces far fewer recognition… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.Hindawi-Books-dataset
Dataset Card for "Hindawi Books Dataset"
Hindawi Books Dataset is a large collection of more than 3000 books written in Modern Standard Arabic.
Dataset Description
Hindawi Books Dataset offers a rich and diverse collection of literary works, covering various topics and genres, all written in Modern Standard Arabic. The dataset includes information about each book, such as the title, author name, book abstract, and a link to access the complete text online. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/alielfilali01/Hindawi-Books-dataset.task1650_opus_books_en-fi_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1650_opus_books_en-fi_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1650_opus_books_en-fi_translation.wiki-ru-en-news-booksrp_books-en
Dataset Card for "rp_books-en"
Filtering/cleaning on the 'red pajama books' subset of togethercomputer/Long-Data-Collections
The default config:
Dataset({
features: ['meta', 'text'],
num_rows: 26372
})
token count
default
GPT-4 tiktoken token count:
token_count
count 2.637200e+04
mean 1.009725e+05
std 1.161315e+05
min 3.811000e+03
25% 3.752750e+04
50% 7.757950e+04
75% 1.294130e+05
max 8.687685e+06
Total count: 2662.85 M… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/rp_books-en.rock-pollub-pl-books
ROCK Politechnika Lubelska PL books
A fail-closed Polish academic-book subset extracted from the official ROCK repository of Lublin University of Technology.
Retained books: 8
Text characters: 3,380,937
Tokens: 1,177,858 (cl100k_base proxy)
Author coverage: 100.0%
License: CC BY-SA 4.0, confirmed for every retained item and matched PDF bitstream
Source period: 2023-2026
The acquisition target was 20 books, but only eight passed the conservative per-file rights gate. The other… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/rock-pollub-pl-books.books_datasetAzerbaijani Books Dataset
Description
This dataset contains 2800 books on different topics in Azerbaijani language. It was created in 2024 and contains 7.8 million sentences.
The books were divided into sentences and pre-filtered.
The dataset included only those sentences where the percentage of letters was at least 80% of the total number of characters.
The sequence of sentences is the same as in books.
Format
The dataset is provided in comma-separated values (CSV) format. Each article is… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/books_dataset.thai-tnhc2-books
Thai TNHC2 Books
This dataset collect all books from TNHC2 corpus.
We clean the dataset to use text to pretraining model and nlp task.
All books: 353 books
License: CC-0
TNHC2 Dataset (Original) have many a lots of details (chapter, author's detail and more). The dataset is clean to pretraining model and nlp task.
TNHC2 coepus is a Thai old books corpus that all books are copyright expired in Thai law (50 years after the author's death).
TNHC2 Dataset (Original):… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-tnhc2-books.whole-books
Whole Books
Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book.
Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large
majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below;
nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are
documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.thai-it-books
Thai IT books
This dataset collects Thai IT books that are the open access books.
license: cc-by-3.0
LoC-PD-Books-preprocessed
LoC-PD-Books: preprocessed
This is the storytracer/LoC-PD-Books dataset with the following preprocessing steps:
apply clean-text package keeping casing and newlines
drop OCR garbled text in first few lines of each example
fix (most) 'hard' newlines w/ regex similar to gutenberg clean
'grade' first 512 tokens of each book with this quantized model; keep examples from labels clean (all) and mild gibberish w/ score 0.9 or higher
task1647_opus_books_en-pt_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1647_opus_books_en-pt_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1647_opus_books_en-pt_translation.Electrical_BooksDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/SUSHANT283/Electrical_Books.crh_booksTheArabicPile_Books
The Arabic Pile
Introduction:
The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Books.task1652_opus_books_ca-en_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1652_opus_books_ca-en_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1652_opus_books_ca-en_translation.Ukrainian-CulturalHeritage-Books
🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦
Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain.
Dataset summary
The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources.
Curation method
The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.task1649_opus_books_en-no_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1649_opus_books_en-no_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1649_opus_books_en-no_translation.task1651_opus_books_en-es__translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1651_opus_books_en-es__translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1651_opus_books_en-es__translation.books-intent-dataset
Books intent classification dataset
A prompt intent classification dataset built from titles, author names and categories (subjects) contained in Project Gutenberg.
Its main purpose is to finetune small language models on intent classification task.
Dataset Details
Dataset Description
Curated by: Empathy.co
Shared by: Project Gutenberg
Language(s) (NLP): English
License: CC0 1.0 Public‑Domain Dedication
Dataset Sources
Project Gutenberg.… See the full description on the dataset page: https://huggingface.co/datasets/empathyai/books-intent-dataset.airship-books-antique
pszemraj/airship-books-antique
Digitized books (via a VLM) on airships from the survivor library
task1648_opus_books_en-sv_translation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1648_opus_books_en-sv_translation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1648_opus_books_en-sv_translation.Books-General-Linux
Linux Books Dataset
Dataset Description
The Linux Books Dataset is a curated text dataset derived from Linux-related books and learning materials. It focuses on Linux system administration, cybersecurity, networking, shell scripting, and operating system fundamentals.The dataset is designed to support training and evaluation of NLP models for technical domains, especially cybersecurity-aware language models and Linux-focused assistants.
This dataset is suitable for both… See the full description on the dataset page: https://huggingface.co/datasets/DexopT/Books-General-Linux.spanish_books
Spanish Books
Dataset Summary
Dataset of books in Spanish crawled from web and torrents.
Preprocessing
Preprocessing performed by spanish_nlp.
Licensing Information
The dataset is available under the Creative Commons Attribution-ShareAlike License (CC BY-SA 4.0).
Some books may be subject to copyright. Use for academic purposes only.
Citation Information
@misc{ortiz2022esbooks,
title={Crawled Spanish Books},
author={Jorge… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/spanish_books.books-gutenberg-project-pt-br
Gutenberg Project TokenWeaver CPT 2048 - Unchunked
This dataset contains full-document rows reconstructed from
costadev00/gutenberg-project-tokenweaver-cpt-2048.
The source dataset mixes reconstructed chunk sequences and singleton chunk rows.
For this unchunked release, rows were grouped by metadata.id; when duplicated
singleton rows were present for the same document, the reconstruction kept the
series with the largest chunk_total. Text was joined with inferred text
overlap.… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/books-gutenberg-project-pt-br.
