CoolFace
Datasetpublic

nikolina-p/mini_gutenberg_splits

Dataset Card for Mini Project Gutenberg Dataset This dataset is a mini subset of the dataset nikolina-p/gutenberg_clean_en, created for learning, testing streaming datasets, and quick downloading and manipulation. It is made from the first 24 books, which are randomly split into 39 shards, mirroring the structure of the original dataset. The text of the books is randomly split into small chunks, allowing users to experiment with dataset operations on a smaller scale. This… See the full description on the dataset page: https://huggingface.co/datasets/nikolina-p/mini_gutenberg_splits.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes18downloads
../
fileshard-000.parquet96 KBdownload
fileshard-001.parquet135 KBdownload
fileshard-002.parquet124 KBdownload
fileshard-003.parquet41 KBdownload
fileshard-004.parquet14 KBdownload
fileshard-005.parquet175 KBdownload
fileshard-006.parquet272 KBdownload
fileshard-007.parquet354 KBdownload
fileshard-008.parquet217 KBdownload
fileshard-009.parquet220 KBdownload
fileshard-010.parquet31 KBdownload
fileshard-011.parquet48 KBdownload
fileshard-012.parquet66 KBdownload
fileshard-013.parquet104 KBdownload
fileshard-014.parquet159 KBdownload
fileshard-015.parquet43 KBdownload
fileshard-016.parquet79 KBdownload
fileshard-017.parquet135 KBdownload
fileshard-018.parquet145 KBdownload
fileshard-019.parquet192 KBdownload
fileshard-020.parquet63 KBdownload
fileshard-021.parquet104 KBdownload
fileshard-022.parquet148 KBdownload
fileshard-023.parquet105 KBdownload
fileshard-024.parquet116 KBdownload
fileshard-025.parquet272 KBdownload
fileshard-026.parquet314 KBdownload
fileshard-027.parquet34 KBdownload
fileshard-028.parquet97 KBdownload
fileshard-029.parquet156 KBdownload
fileshard-030.parquet82 KBdownload
fileshard-031.parquet127 KBdownload
fileshard-032.parquet132 KBdownload

nikolina-p/mini_gutenberg_splits · main · files are served by the source, never re-hosted here