datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pre_1929_books_filtered
Pre-1929 Books
Description
Books published in the US before 1929 passed into the public domain on January 1, 2024.
We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929.
The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries.
The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.pre_1929_books
Pre-1929 Books
Description
Books published in the US before 1929 passed into the public domain on January 1, 2024.
We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929.
The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries.
The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books.booksum-complete-cleaned
Description:
This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization
.
This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.BookSummary
Polish BookSummary
Polish BookSummary is a long-document factual-consistency classification dataset constructed from a publicly available collection of English-language book plot summaries.
Each example consists of a short claim followed by a Polish plot summary. The task is a binary classification problem in which the model must determine whether the claim is factually consistent with the information contained in the document.
Polish BookSummary is part of the LongContext… See the full description on the dataset page: https://huggingface.co/datasets/mmichall/BookSummary.Comprehensive-Antiquarian-and-Rare-Books-Archive
Comprehensive Antiquarian & Rare Books Archive
Dataset Description
This dataset contains pristine, commerce-free bibliographical metadata extracted from the Govi Rare Books Archive. It is engineered to provide high-fidelity, structured historical data for Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines. By supplying ground-truth bibliographical metadata, this repository aims to reduce AI hallucinations and improve semantic reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Govi-Rare-Books-Archive/Comprehensive-Antiquarian-and-Rare-Books-Archive.book_snippets_asrTiny-Open-Domain-BooksA tiny example dataset consisting of four books dedicated to the open domain in JSONL format:
Alice in Wonderland - Lewis Caroll
Dracula - Bram Stoker
The Wonderful Wizard of Oz - L. Frank Baum
The Count of Monte Cristo - Alexandre Dumas & Auguste Maquet
All works are open domain, thus this dataset is also dedicated to the open domain.
The dataset has been made to have extremely long context lengths, ideally as close to 2048 at possible without cutting off chunks in strange places. Each… See the full description on the dataset page: https://huggingface.co/datasets/Blackroot/Tiny-Open-Domain-Books.clean-PD-16000-books3
📚 clean-PD-16000-books3
A treasure trove of ~16,000 high-quality, public domain books in English language — nicely cleaned, with rich metadata, and ready for language modeling.
✨ What Makes This Dataset Special?
This isn’t just another dump of dusty old text files.
clean-PD-16000-books3 is the result of a rigorous cleaning and curation process applied to a large collection of public domain literature, including:
✅ Readable prose — paragraphized prose, without unnatural… See the full description on the dataset page: https://huggingface.co/datasets/next-token/clean-PD-16000-books3.books
Books
The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books.
This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling.
Details
The dataset contains 4 columns:
title: The tilte of the book.
author: The author of the book.
category: The… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/books.EvolvingWorld-Books-Gutenberg
Dataset Card for EvolvingWorld-Books-Gutenberg
Dataset Summary
This dataset provides the 57 source book texts used for EvolvingWorld data construction. The books were selected from Goodreads' Best Books Ever list and obtained from Project Gutenberg.
Dataset Structure
Each row is a JSON object with the following fields:
title: the title of the book
author: the author of the book
content: the full text content used for data construction
Example:
{… See the full description on the dataset page: https://huggingface.co/datasets/zongqing0068/EvolvingWorld-Books-Gutenberg.Book_Summary_Chinese
中文图书总结数据集
每个样本包含:
图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。
tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.zer0-books
scraped_at_utc: 2026-03-24T12:45:50.800626+00:00,
source: https://www.collectiveinkbooks.com/zer0-books/our-books/all-books/&s=0,
listing_total_books: 383,
listing_total_pages: 16,
scraped_books: 383,
books_with_reviews: 341,
books_without_reviews: 42.
This dataset contains the list of all the books that have been, as of 17.04.26, published by a British independent philosophical publisher zer0 books founded by Mark Fisher. No copyright infringement had been made, and all the content is… See the full description on the dataset page: https://huggingface.co/datasets/mstyslavity/zer0-books.zer0-books
scraped_at_utc: 2026-03-24T12:45:50.800626+00:00,
source: https://www.collectiveinkbooks.com/zer0-books/our-books/all-books/&s=0,
listing_total_books: 383,
listing_total_pages: 16,
scraped_books: 383,
books_with_reviews: 341,
books_without_reviews: 42.
This dataset contains the list of all the books that have been, as of 17.04.26, published by a British independent philosophical publisher zer0 books founded by Mark Fisher. No copyright infringement had been made, and all the content is… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/zer0-books.booksum-zhbooksum数据集,谷歌翻译成中文。
任务:将一本书的某个章节总结为几句话。
源数据来自 togethercomputer/Long-Data-Collections
Long-Data-Collections-booksum-binidxFine-tune Data
BookSum:
BookSum is a dataset for long context summarization. It includes a vast collection of books from various genres, and the task is to generate a coherent and concise summary given a long context from the book. This dataset is designed to test and train models on their ability to understand and summarize long, complex narratives.
to convert to binidx format.
pile_books3booksum_deA german translation for the booksum dataset.
Extracted from seedboxventures/multitask_german_examples_32k.
Translation created by seedbox ai for KafkaLM ❤️.
Available for finetuning in hiyouga/LLaMA-Factory.
apollo_english_books_translated_to_dutch_with_geminiflash15
Data description
Translation of the English medical books that are part of the Apollo corpus, using the LLM Gemini Flash 1.5
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
beige-booksafricans-history-books-qa-testturkish-classic-books-qa58_books_religious_textThis dataset was made from 58 books on religion. The books were only processed by cleaning scripts and still contain elemets such as page nubers, table of contents, glossaries and other unwanted data.
This was created to establish a tuning workflow with the intention of improving the data later.
africans-history-books-qa-evalbooks
Books
The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books.
This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling.
Details
The dataset contains 4 columns:
title: The tilte of the book.
author: The author of the book.
category: The… See the full description on the dataset page: https://huggingface.co/datasets/DhruvExploring/books.defendable-pain-books-and-records-receipts-v0.1
Books and Records Receipts · DefendableLedger Sample
"the chain" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 30 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-books-and-records-receipts-v0.1.richard-amir-nasser-ai-research-books
Richard Amir Nasser AI Research Books Evidence Map
This public-safe dataset card maps the strongest AI, search, roofing documentation, and research-system books connected to Richard Amir Nasser.
It is intended as an evidence map, not a ranking claim, credential claim, legal claim, insurance claim, or certification.
Person Frame
Richard Amir Nasser is framed here as a founder, author, and research systems developer focused on inspection-first roofing… See the full description on the dataset page: https://huggingface.co/datasets/InspectorRoofing/richard-amir-nasser-ai-research-books.Orion-Light-Novels-Roleplay-Logs-Books-Oh-My-duplicate-turns-removedbooks-ner-dataset
Books Named Entity Recognition (NER) Dataset
A lightweight Named‑Entity‑Recognition (NER) corpus built from titles and author names contained in Project Gutenberg’s public catalogues. It is intended for training or benchmarking entity extractors such as Gliner on bibliographic metadata.
1 Provenance
This dataset provenance originates from Project Gutenberg's public catalogue.
2 Quick facts
Records (total)
434 925
Train split
391 432 queries… See the full description on the dataset page: https://huggingface.co/datasets/empathyai/books-ner-dataset.Books_of_the_period_of_representation_of_Zaznobin_V_M
