CoolFace
Datasetpublic

nikolina-p/mini_gutenberg_splits

Dataset Card for Mini Project Gutenberg Dataset This dataset is a mini subset of the dataset nikolina-p/gutenberg_clean_en, created for learning, testing streaming datasets, and quick downloading and manipulation. It is made from the first 24 books, which are randomly split into 39 shards, mirroring the structure of the original dataset. The text of the books is randomly split into small chunks, allowing users to experiment with dataset operations on a smaller scale. This… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/mini_gutenberg_splits.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes18downloads

nikolina-p/mini_gutenberg_splits · main · files are served by the source, never re-hosted here