CoolFace
Datasetpublic

gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk8-data

GSA-PT-Llama-3.2-1B-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk8 — model trained on this dataset

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes254downloads

gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk8-data · main · files are served by the source, never re-hosted here