CoolFace
Apppublic

alphonse86/AI-resources

sourceHugging Facemitupdated 5mo agoView on Hugging Face
0likes
App README

AI-resources

A List of Foundational AI / ML / LLM Papers

Encoder-Decoder Architecture

Attention Is All You Need - Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin link

Parallel Distribution Processing

Microprocessor Architecture


Major Themes of Modern LLM Transformers:

Attention

Deep Learning (NN)

Hardware/GPU alignments


Rectified Linear Units Improve Restricted Boltzmann Machines - Vinod Nair, Geoffrey E. Hinton link ZeRo: Memory Optimizations Toward Training Trillion Parameter Models - Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He Beyond Regression: New Tools For Prediction and Analysis in the Behavioral Sciences - Paul John Webos ImageNet Classification with Dep Convolutional Neural Networks - Alex Krizhevsky, Ilya Sutskever, Geoffrey E. Hinton link Taylor Expansion of the Accumulated Rounding Erro - Seppo Linnainmaa link Learning to Translate in Real-time with Neural Machine Translation - Jiatao Gu, Graham Neubig, Kyunghyun Cho, Victor O.k. Li link Effective Approaches to Attention-base Neural Machine Translation - Minh-Thang Luong, Hieu Pham, Christopher D. Manning link Neural Machine Translation with Supervised Attention - Lemao Liu, Masao Utiyama, Andrew Finch, Eiichiro Sumita link A Structured Self-Attentive Sentence Embedding - Zhouhan Lin, Minwei Feng, Cicero Nogueria dos Santos, Mo Yu, Bing Xiang, Bown Zhou, Yoshua Bengio link FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness - Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re link Learning Representations by back-propagating error - David E. Rumelhard, Geoffrey E. Hinton, Ronald J. Williams link Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Paralleism - Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catazaro link End-to-end Memory Networks - Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus link Memory Networks - Jason Weston, Sumit Chopra, Antoine Bordes link Neural Machine Translation by Jointly Learning to Align and Translate - Dzmitry Bahdanau, KyungHyun Cho, Yoshua Bengio link Approximation by Superpositions of a Sigmoidal Function - G. Cybenko link