CoolFace
Datasetpublic

chanind/pile-uncopyrighted-gemma-1024-abbrv-2B

Pre-tokenized dataset of the first 10 million lines of monology/pile-uncopyrighted without any concatenated lines, tokenized for Gemma-2 using SAELens. This dataset has 1024 context size and about 2.5B tokens.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes670downloads
Dataset Card

Pre-tokenized dataset of the first 10 million lines of monology/pile-uncopyrighted without any concatenated lines, tokenized for Gemma-2 using SAELens.

This dataset has 1024 context size and about 2.5B tokens.