chanind/pile-uncopyrighted-gemma-1024-abbrv-2B
Pre-tokenized dataset of the first 10 million lines of monology/pile-uncopyrighted without any concatenated lines, tokenized for Gemma-2 using SAELens. This dataset has 1024 context size and about 2.5B tokens.
0681
Update README.md
Add sae_lens metadata
Upload dataset (part 00001-of-00002)
Upload dataset (part 00000-of-00002)
initial commit
