CoolFace
Datasetpublic

Navanjana/Gutenberg_books

Gutenberg Books Dataset Dataset Description This dataset contains 97,646,390 paragraphs extracted from 74,329 English-language books sourced from Project Gutenberg, a digital library of public domain works. The total size of the dataset is 34GB, making it a substantial resource for natural language processing (NLP) research and applications. The texts have been cleaned to remove Project Gutenberg's standard headers and footers, ensuring that only the core content… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/Gutenberg_books.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
5likes79downloads

Navanjana/Gutenberg_books · main · files are served by the source, never re-hosted here