alphonse86/AI-resources
AI-resources
A List of Foundational AI / ML / LLM Papers
Encoder-Decoder Architecture
Attention Is All You Need - Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin link
Parallel Distribution Processing
Microprocessor Architecture
Major Themes of Modern LLM Transformers:
Attention
Deep Learning (NN)
Hardware/GPU alignments
Rectified Linear Units Improve Restricted Boltzmann Machines - Vinod Nair, Geoffrey E. Hinton link ZeRo: Memory Optimizations Toward Training Trillion Parameter Models - Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He Beyond Regression: New Tools For Prediction and Analysis in the Behavioral Sciences - Paul John Webos ImageNet Classification with Dep Convolutional Neural Networks - Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton link Taylor Expansion of the Accumulated Rounding Erro - Seppo Linnainmaa link Learning to Translate in Real-time with Neural Machine Translation - Jiatao Gu, Graham Neubig, Kyunghyun Cho, Victor O.k. Li link Effective Approaches to Attention-base Neural Machine Translation - Minh-Thang Luong, Hieu Pham, Christopher D. Manning link Neural Machine Translation with Supervised Attention - Lemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita link A Structured Self-Attentive Sentence Embedding - Zhouhan Lin, Minwei Feng, Cicero Nogueria dos Santos, Mo Yu, Bing Xiang, Bown Zhou, Yoshua Bengio link FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness - Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re link Learning Representations by back-propagating error - David E. Rumelhard, Geoffrey E. Hinton, Ronald J. Williams link Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Paralleism - Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catazaro link End-to-end Memory Networks - Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus link Memory Networks - Jason Weston, Sumit Chopra, Antoine Bordes link Neural Machine Translation by Jointly Learning to Align and Translate - Dzmitry Bahdanau, KyungHyun Cho, Yoshua Bengio link Approximation by Superpositions of a Sigmoidal Function - G. Cybenko link
