CoolFace
Datasetpublic

chanind/pile-uncopyrighted-gemma-1024-abbrv-2B

Pre-tokenized dataset of the first 10 million lines of monology/pile-uncopyrighted without any concatenated lines, tokenized for Gemma-2 using SAELens. This dataset has 1024 context size and about 2.5B tokens.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes681downloads
5 commits on main
1a48ea51y ago

Update README.md

chanind
4beb0131y ago

Add sae_lens metadata

chanind
599768c1y ago

Upload dataset (part 00001-of-00002)

chanind
0d227151y ago

Upload dataset (part 00000-of-00002)

chanind
25c8fdd1y ago

initial commit

chanind