ysngkil/whole-books
Whole Books Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book. Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below; nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face