CoolFace
Datasetpublic

LisaMegaWatts/bookcorpus-gutenberg-classics

BookCorpus + Gutenberg Classics Training Corpus Large-scale training corpus combining BookCorpus fiction, Project Gutenberg 19th-century literature (PG-19), and curated classical philosophy texts. Cleaned, deduplicated, and organized into curriculum phases for character-level language model training. Dataset Description This corpus is the primary training dataset for the Julia SLM project, combining three major text sources into a unified, cleaned training set… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/bookcorpus-gutenberg-classics.

sourceHugging Faceapache-2.0updated 7mo agoView on Hugging Face
0likes115downloads

LisaMegaWatts/bookcorpus-gutenberg-classics · main · files are served by the source, never re-hosted here