CoolFace
Modelpublic

bailin28/gla-1B-100B

sourceHugging Facemitupdated 3y agoView on Hugging Face
2likes29downloads
Model Card

This checkpoint of the 1.3B GLA model used in the paper Gated Linear Attention. The model is trained with 100B tokens from the SlimPajama dataset tokenized with Llama2 tokenizer.

See the model and loading script in this repo.