CoolFace
Datasetpublic

Navanjana/Gutenberg_books

Gutenberg Books Dataset Dataset Description This dataset contains 97,646,390 paragraphs extracted from 74,329 English-language books sourced from Project Gutenberg, a digital library of public domain works. The total size of the dataset is 34GB, making it a substantial resource for natural language processing (NLP) research and applications. The texts have been cleaned to remove Project Gutenberg's standard headers and footers, ensuring that only the core content… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/Gutenberg_books.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
5likes68downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
Navanjana/Gutenberg_books · CoolFace