CoolFace
Datasetpublic

nikolina-p/mini_gutenberg_splits

Dataset Card for Mini Project Gutenberg Dataset This dataset is a mini subset of the dataset nikolina-p/gutenberg_clean_en, created for learning, testing streaming datasets, and quick downloading and manipulation. It is made from the first 24 books, which are randomly split into 39 shards, mirroring the structure of the original dataset. The text of the books is randomly split into small chunks, allowing users to experiment with dataset operations on a smaller scale. This… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/mini_gutenberg_splits.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes18downloads
7 commits on main
2b9b1a71y ago

Update README.md

nikolina-p
afd9d891y ago

Update README.md

nikolina-p
966a0511y ago

Update README.md

nikolina-p
fb0e6fd1y ago

Update README.md

nikolina-p
e4a39881y ago

Create README.md

nikolina-p
12e35fd1y ago

Initial upload of dataset

nikolina-p
0749a401y ago

initial commit

nikolina-p