CoolFace
Datasetpublic

NNEngine/Gutenberg-Clean

๐Ÿ“š TinyWay-Gutenberg-Clean (Compressed Shards) A large-scale, high-quality English text dataset derived from Project Gutenberg. The corpus has been cleaned, normalized, deduplicated, segmented into fixed-length samples, and stored as compressed JSONL shards for efficient large-scale language model training. This dataset is intended for pretraining and experimentation with small and medium language models such as TinyWay, tokenizer training, and large-scale NLP research.โ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Gutenberg-Clean.

sourceHugging Facemitupdated 8mo agoView on Hugging Face
0likes54downloads
3 commits on main
a41180b8mo ago

Update README.md

NNEngine
422e21d8mo ago

Initial release: Sentiment Analysis Dataset

NNEngine
4e6d7668mo ago

initial commit

NNEngine