datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining_v1-omega_booksarabic-books
Arabic Books
Dataset Summary
The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer.
This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.kb-books
open-rdl-books
Dataset Description
Language
dan, dansk, Danish
License
Public Domain, cc0-1.0
Dataset Summary
Documents from the Royal Danish Library published between 1750 and 1930.
The dataset has each page of each document in image and text format. The text was extracted with OCR.
The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible.… See the full description on the dataset page: https://huggingface.co/datasets/chcaa/kb-books.noor-platform-booksopus_books
Dataset Card for OPUS Books
Dataset Summary
This is a collection of copyright free books aligned by Andras Farkas, which are available from http://www.farkastranslations.com/bilingual_books.php
Note that the texts are rather dated due to copyright issues and that some of them are manually reviewed (check the meta-data at the top of the corpus files in XML). The source is multilingually aligned, which is available from http://www.farkastranslations.com/bilingual_books.php.… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/opus_books.institutional-books-hl
📚 Institutional Books: Harvard Library
Institutional Books is a growing corpus of public domain books. This release is comprised of 983,004 public domain books digitized as part of Harvard Library's participation in the Google Books project and refined by the Institutional Data Initiative. Use of this data is governed by the IDI Terms of Use for Early-Access.
983K books, published largely in the 19th and 20th centuries
242B o200k_base tokens
386M pages of text, available in… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl.vi-books_tve-4u.orgTODOs
dtv-ebook.com bóc tách và convert dtv ebooks 3543 files | 4.5G (without PDF)
dtv_categories_whitelist.jsonl cho vào ../zinz/30-vien_book_epub_van-hoc/
phần còn lại để vào ../zinz/40-vi_character_entertain/
xếp dữ liệu theo từng categories
tve-4u.org bóc tách và convert thành text ~11k files | 7.2G (without PDF)
download_links_with_post_url liệt kê toàn bộ download links để download ebooks và truy xuất ngược về post có chưa download link tương ứng
posts_with_content là các… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/vi-books_tve-4u.org.institutional-books-hl-visual-elements
📚 Institutional Books: Harvard Library — Visual Elements
22 million visual elements extracted from the volumes that comprise the Institutional Books: Harvard Library dataset.
22,622,060 visual elements extracted from 983,004 volumes
766,992,447 o200k_base tokens in AI-generated captions
6 high-level classes of visual elements organized in splits
5 processing steps: Detection, Classification, Deduplication, Captioning, and Rotation
The Institutional Data Initiative at Harvard… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-visual-elements.vi-books_extractionll books_text_final/ | wc -l
=> 13968 text
After filter (not vi, too small)
wc -l vi-books.jsonl
=> 12832 text
=> !!! Cần lọc truyện từ dtv ra !!!
xzcat dtv_ebooks_details.jsonl.xz | wc -l
# => 13486
ll books_text_final/ | grep dtv_ebooks_ | wc -l
# => 3535 từ dtv
lấy các whitelist categories tại dtv_categories_whitelist.jsonl
lần theo từng ebook một trong dtv_ebooks_details.jsonl.xz trường {"cat": "Self Help - Khởi nghiệp", "cat_url":… See the full description on the dataset page: https://huggingface.co/datasets/tiendung/vi-books_extraction.US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes.
I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.institutional-books-hl-enriched-text
📚 Institutional Books: Harvard Library — Enriched Text
Institutional Books is a growing corpus of public domain books.
This release (IB-HL-ET) is a version of the text present in the Institutional Books: Harvard Library
(IB-HL) dataset
that has been further processed, filtered and optimized for computational access and model training.
This includes:
983K books, published largely in the 19th and 20th centuries
217B o200k_base tokens
7B sentences in 250 languages, grouped into… See the full description on the dataset page: https://huggingface.co/datasets/institutional/institutional-books-hl-enriched-text.booksum
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization
Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev
Introduction
The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases.
While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.MUSE-Books
MUSE-Books
MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-Books.1k-books-corpus
Public Domain German Books (1800–1879)
A page-level dataset of 1,111 German books (299,610 pages) printed between
1800 and 1879, digitised by the Münchener DigitalisierungsZentrum (MDZ)
of the Bayerische Staatsbibliothek. Every page comes with its full-resolution
scan, the hOCR output, plain text extracted from that hOCR, and the IIIF
Image API description of the scan.
All images, info.jsons and hOCR were downloaded from the official APIs.
Dataset Structure
One… See the full description on the dataset page: https://huggingface.co/datasets/histde/1k-books-corpus.vi-books_dtv-ebook.com
git lfs track "*.epub"
git lfs track "*.mobi"
git lfs track "*.prc"
git lfs track "*.txt"
git lfs track "*.json"
git lfs track "*.jsonl"
git lfs track "*.rar"
git lfs track "*.zip"
git lfs track "*.doc"
git lfs track "*.docx"
git lfs track "*.pdf"
git lfs track "*.azw3"
du -sh download/* > _size_downloads.txt
French-PD-Books
🇫🇷 French Public Domain Books 🇫🇷
French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.arabic-synthetic-scanned-booksocr_arabic_books
Arabic OCR Books Dataset (ocr_arabic_books)
This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions.
📚 Master Book Inventory
Total running pages in repository: 181,427
#
Book Name (English)
Book Name (Arabic)
Subset / Config Name
Page Count
Image Index Range
1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.green-books-thumbnails
African American Travel Guides: Listing Thumbnails
113,053 pre-cropped thumbnail images — one per listing — from 50 volumes of mid-20th-century African American travel guides (1930–1966). Each image is a small snippet of a scanned directory page cropped down to a single business or lodging listing: the exact region a traveler would have read.
This is the image companion to the structured-listings dataset hadro/green-books-travel-guides. That dataset holds the transcribed text of… See the full description on the dataset page: https://huggingface.co/datasets/hadro/green-books-thumbnails.books_and_conversationsBooksbooksummaries_cleanedsynth_shamela_ocr_arabic_books
Synthetic Arabic Books Dataset
Structured book pages rendered dynamically with style, font, and degradation variations.
pre_1929_books_filtered
Pre-1929 Books
Description
Books published in the US before 1929 passed into the public domain on January 1, 2024.
We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929.
The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries.
The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.Long-Data-Collections-Pretrain-Without-Books
Dataset Card for "Long-Data-Collections-Pretrain-Without-Books"
Paraquet version of the pretrain split of togethercomputer/Long-Data-Collections WITHOUT books
Statistics (in # of characters): total_len: 236088622215, average_len: 25159.041601590307
bl_books_flickr_fullSpanish-PD-Books
🇪🇸 Spanish Public Domain Books 🇪🇸
Spanish-Public Domain-Newspapers or Spanish-PD-Newspapers is a large collection aiming to aggregate all Spanish monographies in the public domain. As of March 2024, with Spanish-PD-Newspapers, it is the biggest Spanish open corpus.
Dataset summary
The collection contains 302,640 individual texts making up 13.9 billion words recovered from multiple sources, including Spanish leading cultural heritage institution Biblioteca Digitale… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Spanish-PD-Books.sanskrit_books_collectionLoC-PD-Books
Library of Congress Public Domain Books (English)
This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic.
Curation method
The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books.
Dataset summary
The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.pg_books-tokenized-bos-eos-chunked-65536
Dataset Card for "pg_books-tokenized-bos-eos-chunked-65536"
The pg19 dataset tokenized under LLaMA into 64k chunks, bookended with BOS and EOS
