CoolFace
Datasetpublic

ysngkil/whole-books

Whole Books Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book. Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below; nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.

sourceHugging Faceotherupdated 22d agoView on Hugging Face
0likes65downloads
settings

This repository belongs to ysngkil on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namewhole-books
visibilitypublic
licenceother
gatedno
ownerysngkil
Account settings
ysngkil/whole-books · CoolFace