CoolFace
Datasetpublic

gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk8-data

GSA-PT-Qwen2-7B-Instruct-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes58downloads
filedata-00000-of-00015.arrow562.2 MBdownload
filedata-00001-of-00015.arrow411.6 MBdownload
filedata-00002-of-00015.arrow488.5 MBdownload
filedata-00003-of-00015.arrow481.0 MBdownload
filedata-00004-of-00015.arrow485.0 MBdownload
filedata-00005-of-00015.arrow494.9 MBdownload
filedata-00006-of-00015.arrow474.0 MBdownload
filedata-00007-of-00015.arrow466.1 MBdownload
filedata-00008-of-00015.arrow460.8 MBdownload
filedata-00009-of-00015.arrow463.8 MBdownload
filedata-00010-of-00015.arrow456.5 MBdownload
filedata-00011-of-00015.arrow457.6 MBdownload
filedata-00012-of-00015.arrow459.4 MBdownload
filedata-00013-of-00015.arrow443.7 MBdownload
filedata-00014-of-00015.arrow463.5 MBdownload

gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk8-data · main · files are served by the source, never re-hosted here