CoolFace
Datasetpublic

gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk16-data

GSA-PT-Llama-3.2-1B-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk16 — model trained on this dataset

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes142downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk16-data · CoolFace