CoolFace
Datasetpublic

ysngkil/whole-books

Whole Books Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book. Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below; nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.

sourceHugging Faceotherupdated 22d agoView on Hugging Face
0likes65downloads

ysngkil/whole-books · main · files are served by the source, never re-hosted here