CoolFace
Datasetpublic

Navanjana/Gutenberg_books

Gutenberg Books Dataset Dataset Description This dataset contains 97,646,390 paragraphs extracted from 74,329 English-language books sourced from Project Gutenberg, a digital library of public domain works. The total size of the dataset is 34GB, making it a substantial resource for natural language processing (NLP) research and applications. The texts have been cleaned to remove Project Gutenberg's standard headers and footers, ensuring that only the core content… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/Gutenberg_books.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
5likes68downloads
7 commits on main
4cc91731y ago

Update README.md

Navanjana
27a352c1y ago

Update README.md

Navanjana
60b28391y ago

Update README.md

Navanjana
4beb8ab1y ago

Update README.md

Navanjana
afa220a1y ago

Update README.md

Navanjana
3215d241y ago

Upload file.csv with huggingface_hub

Navanjana
3e26bfe1y ago

initial commit

Navanjana