CoolFace
Datasetpublic

gist-sparse-attention/GSA-PT-Llama-3.2-1B-chunk16-data

GSA-PT-Llama-3.2-1B-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk16 — model trained on this dataset

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes142downloads
4 commits on main
42e8e106mo ago

Upload README.md with huggingface_hub

yuzhenm
8f35f446mo ago

Upload folder using huggingface_hub

yuzhenm
0ec482e6mo ago

Upload folder using huggingface_hub

yuzhenm
d1d79706mo ago

initial commit

yuzhenm