CoolFace
Datasetpublic

GulkoA/TinyStories-tokenized-Llama-3.2

TinyStories dataset tokenized with Llama-3.2 Useful for accelerated training and testing of sparse autoencoders Context window: 128, not shuffled For first layer activations cache with Llama-3.2-1B, see GulkoA/TinyStories-Llama-3.2-1B-cache

sourceHugging Facecdla-sharing-1.0updated 2y agoView on Hugging Face
1likes127downloads

GulkoA/TinyStories-tokenized-Llama-3.2 · main · files are served by the source, never re-hosted here