aecetin/Qwen-2.5-7B-PolySurgery
0386
🥋 Alibaba Qwen-2.5-7B-PolySurgery
Subtitle: Eliminating 4.26 Billion Parameters (-74.78% SwiGLU Weights) from Qwen-2.5-7B via Closed-Form Chebyshev Polynomial Tensor Surgery. Author: Dr. A. Emre ÇETİN (Computational Systems and Cognitive Architectures, Izmir, Turkey) Official Paper: Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks Patent Base: U.S. Patent Application No.64/149,540&64/148,668GitHub Repository: github.com/aemre-cetin/idempotent-poly
💡 What is Qwen-2.5-7B-PolySurgery?
Alibaba's Qwen-2.5-7B has an extraordinarily wide intermediate feed-forward dimension: $d=3584$ with $d_{ff}=18944$. This results in 203.69 Million parameters per layer in standard SwiGLU across 28 layers, consuming $5.70$ Billion parameters (75% of the total model).
Through Orthogonal Polynomial Surgery, this massive bottleneck is converted into a 4-term orthogonal tensor ($K=3$), reducing per-layer FFN parameters to 51.38 Million (-74.78% reduction).
📊 Benchmark Results
📜 Citation
@article{cetin2026orthogonalpoly,
title={Hardware-Accelerated Orthogonal Polynomial Tensor Operators, Zero-Backpropagation Closed-Form Algebraic Solvers, and In-Situ Weight Surgery for Deep Neural Networks},
author={Çetin, A. Emre},
journal={arXiv preprint arXiv:2609.xxxxx},
year={2026}
}