NNEngine/Gutenberg-Clean
๐ TinyWay-Gutenberg-Clean (Compressed Shards) A large-scale, high-quality English text dataset derived from Project Gutenberg. The corpus has been cleaned, normalized, deduplicated, segmented into fixed-length samples, and stored as compressed JSONL shards for efficient large-scale language model training. This dataset is intended for pretraining and experimentation with small and medium language models such as TinyWay, tokenizer training, and large-scale NLP research.โฆ See the full description on the dataset page: https://huggingface.co/datasets/NNEngine/Gutenberg-Clean.
054
Update README.md
Initial release: Sentiment Analysis Dataset
initial commit
