CoolFace
Datasetpublic

gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data

GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this… See the full description on the dataset page: https://huggingface.co/datasets/gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data.

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes52downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
gist-sparse-attention/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data · CoolFace