gist-sparse-attention/GSA-FT-Llama-3.2-1B-data
GSA-FT-Llama-3.2-1B-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Llama-3.2-1B. The answers to the questiosn are generated by Deepseek-V3.2 Each sample is formatted with GSA gist tokens for instruction fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-FT-Llama-3.2-1B-chunk8 yuzhenm/GSA-FT-Llama-3.2-1B-chunk16… See the full description on the dataset page: https://huggingface.co/datasets/gist-sparse-attention/GSA-FT-Llama-3.2-1B-data.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face