datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset
repro-incremental-learning-of-sparse-attention-patterns-in-transformers-bundle
Reproduction bundle — Incremental Learning of Sparse Attention Patterns in Transformers
ICML 2026 paper #12503 · OpenReview vSRh1qU5sH · arXiv 2602.19143
Everything needed to re-run the reproduction whose results are recorded in the
Trackio logbook GwendalTsang/repro-incremental-learning-of-sparse-attention-patterns-in-transformers.
Layout
upstream/ the authors' official code, vendored unchanged… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/repro-incremental-learning-of-sparse-attention-patterns-in-transformers-bundle.GSA-PT-Qwen2-7B-Instruct-chunk16-data
GSA-PT-Qwen2-7B-Instruct-chunk16-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset
GSA-PT-Llama-3.2-1B-chunk8-data
GSA-PT-Llama-3.2-1B-chunk8-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk8.
Each sample is tokenized and formatted with GSA gist tokens for continued pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Models
yuzhenm/GSA-PT-Llama-3.2-1B-chunk8 — model trained on this dataset
GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data
GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset
GSA-PT-Llama-3.2-1B-chunk16-data
GSA-PT-Llama-3.2-1B-chunk16-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for continued pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Models
yuzhenm/GSA-PT-Llama-3.2-1B-chunk16 — model trained on this dataset
GSA-FT-Qwen2-7B-Instruct-chunk32-data
GSA-FT-Qwen2-7B-Instruct-chunk32-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset
GSA-PT-Qwen2-7B-Instruct-chunk32-data
GSA-PT-Qwen2-7B-Instruct-chunk32-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset
GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data
GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset
GSA-PT-Qwen2-7B-Instruct-chunk8-data
GSA-PT-Qwen2-7B-Instruct-chunk8-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset
GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset
GSA-PT-Llama-3.2-1B-chunk4-chunk4-data
GSA-PT-Llama-3.2-1B-chunk4-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continued pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Models
yuzhenm/GSA-PT-Llama-3.2-1B-chunk4-chunk4 — model trained on this dataset
GSA-FT-Qwen2-7B-Instruct-chunk16-data
GSA-FT-Qwen2-7B-Instruct-chunk16-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset
GSA-FT-Llama-3.2-1B-data
GSA-FT-Llama-3.2-1B-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Llama-3.2-1B.
The answers to the questiosn are generated by Deepseek-V3.2
Each sample is formatted with GSA gist tokens for instruction fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Models
yuzhenm/GSA-FT-Llama-3.2-1B-chunk8
yuzhenm/GSA-FT-Llama-3.2-1B-chunk16… See the full description on the dataset page: https://huggingface.co/datasets/gist-sparse-attention/GSA-FT-Llama-3.2-1B-data.GSA-FT-Qwen2-7B-Instruct-chunk8-data
GSA-FT-Qwen2-7B-Instruct-chunk8-data
This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8.
Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset
WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B
Sparse Attention for Web Agents — Offline Replay Records
Step-level records from an offline replay study comparing three sparse-attention methods against
full attention on browser-agent trajectories.
Model: Qwen3-VL-30B-A3B-Instruct (48 layers, 128 experts / 8 active, GQA 32:4, head_dim 128)
Hardware: NVIDIA GB10 (sm_121)
Tasks: 50 complete trajectories sampled from WebVoyager + GAIA — 332 steps, of which 190 emit an
element index. Every task was completed successfully by the… See the full description on the dataset page: https://huggingface.co/datasets/shiqihe/WebVoyager-GAIA-SparseAttention-Offline-Qwen3-VL-30B-A3B.
