CoolFace
Datasetpublic

gist-sparse-attention/GSA-FT-Llama-3.2-1B-data

GSA-FT-Llama-3.2-1B-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Llama-3.2-1B. The answers to the questiosn are generated by Deepseek-V3.2 Each sample is formatted with GSA gist tokens for instruction fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-FT-Llama-3.2-1B-chunk8 yuzhenm/GSA-FT-Llama-3.2-1B-chunk16… See the full description on the dataset page: https://huggingface.co/datasets/gist-sparse-attention/GSA-FT-Llama-3.2-1B-data.

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes21downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
gist-sparse-attention/GSA-FT-Llama-3.2-1B-data · CoolFace