CoolFace
Datasetpublic

chanind/openwebtext-gemma

OpenWebTextCorpus tokenized for Gemma This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset. This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it). This dataset was created using SAELens, with the following settings: context_size: 8192… See the full description on the dataset page: https://huggingface.co/datasets/chanind/openwebtext-gemma.

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes779downloads
6 commits on main
155f5242y ago

Update README.md

chanind
5feb9da2y ago

Update README.md

chanind
7650f9f2y ago

Add sae_lens metadata

chanind
38e3c7e2y ago

Upload dataset (part 00001-of-00002)

chanind
ee69fc82y ago

Upload dataset (part 00000-of-00002)

chanind
eb3da952y ago

initial commit

chanind