sparse-attention
GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset
repro-incremental-learning-of-sparse-attention-patterns-in-transformers-bundle
Reproduction bundle — Incremental Learning of Sparse Attention Patterns in Transformers
ICML 2026 paper #12503 · OpenReview vSRh1qU5sH · arXiv 2602.19143
Everything needed to re-run the reproduction whose results are recorded in the
Trackio logbook GwendalTsang/repro-incremental-learning-of-sparse-attention-patterns-in-transformers.
Layout
upstream/ the authors' official code, vendored unchanged… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/repro-incremental-learning-of-sparse-attention-patterns-in-transformers-bundle.GSA-PT-Qwen2-7B-Instruct-chunk16-data
GSA-PT-Qwen2-7B-Instruct-chunk16-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset
GSA-PT-Llama-3.2-1B-chunk8-data
GSA-PT-Llama-3.2-1B-chunk8-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk8.
Each sample is tokenized and formatted with GSA gist tokens for continued pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Models
yuzhenm/GSA-PT-Llama-3.2-1B-chunk8 — model trained on this dataset
GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data
GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset
GSA-PT-Llama-3.2-1B-chunk16-data
GSA-PT-Llama-3.2-1B-chunk16-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for continued pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Models
yuzhenm/GSA-PT-Llama-3.2-1B-chunk16 — model trained on this dataset
repro-tilesparse-sparse-attentionrepro-adasplash-2-faster-differentiable-sparse-attentionrepro-incremental-learning-of-sparse-attention-patterns-in-transformersrepro-tilesparse-arithmetic-intensity-aware-sparse-attention-for-compute-bound-llm-decodingrepro-a-unified-sparse-attention-via-multi-granularity-compressrepro-stochastic-sparse-attention-for-memory-bound-inference
