CoolFace
Modelpublic

SeerAttention/SeerAttention-Decode-Qwen3-8B-AttnGates

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes90downloads
Model Card

SeerAttention-R

This repo contains the decode stage AttnGate weights from paper SeerAttention-R. The current support models are:

Results of Reasoning Tasks

Results of reasoning task with different token budgets. All the results are the averaged pass@1 results with 64 sample per query for AIME, 16 samples for GPQA, and 8 samples for MATH-500.

AIME24

Model2k4k6k8kFull Attention
Qwen3-4B55.4268.7570.9472.5071.25
Qwen3-8B56.5672.2974.2275.0574.48
Qwen3-14B62.2475.7878.0278.6578.91
DeepSeek-R1-Distill-Qwen-14B55.7866.3567.5066.8267.50

AIME25

Model2k4k6k8kFull Attention
Qwen3-4B45.7357.6060.2062.9066.41
Qwen3-8B42.6056.7760.3164.1767.86
Qwen3-14B46.6762.6667.1969.0170.21
DeepSeek-R1-Distill-Qwen-14B38.4447.1952.2550.0550.00

MATH500

Model1k2k4k6kFull Attention
Qwen3-4B84.8092.2093.6093.6093.93
Qwen3-8B82.8291.5394.1794.5394.43
Qwen3-14B85.1393.2094.7794.8095.22
DeepSeek-R1-Distill-Qwen-14B87.6592.1093.0593.1293.30

GPQA Diamond

Model1k2k4k6kFull Attention
Qwen3-4B39.6151.2055.2055.9056.19
Qwen3-8B37.5954.3259.6060.4860.54
Qwen3-14B44.5459.7263.7664.2065.25
DeepSeek-R1-Distill-Qwen-14B51.2656.7956.4157.4857.80