Attention is actually all you need
On brilliance through selective ignorance
2017-12-20 — 2022-08-05
Wherein the Structure of Transformer Stacks and Self‑attention Layers Is Described, the Role in Processing Sequential Data Such as Text Is Examined, and Recent Optimizations Such as FlashAttention Are Noted.
language
machine learning
networks
neural nets
NLP
Bookmarked: A short essay in taking attention absolutely seriously.
1 References
Athey, Tibshirani, and Wager. 2019. “Generalized Random Forests.” Annals of Statistics.
Balestriero, Pesenti, and LeCun. 2021. “Learning in High Dimension Always Amounts to Extrapolation.”
Bayat, Pezeshki, Dohmatob, et al. 2024. “The Pitfalls of Memorization: When Memorization Hurts Generalization.”
Bonnasse-Gahot. 2022. “Interpolation, Extrapolation, and Local Generalization in Common Neural Networks.”
Cao. 2021. “Choose a Transformer: Fourier or Galerkin.” In Advances in Neural Information Processing Systems.
Choy, Gwak, Savarese, et al. 2016. “Universal Correspondence Network.” In Advances in Neural Information Processing Systems 29.
Feldman. 2020. “Does Learning Require Memorization? A Short Tale about a Long Tail.” In Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing. STOC 2020.
Gelman, Hill, and Vehtari. 2021. Regression and Other Stories.
Goddard, Smith, Ngampruetikorn, et al. 2025. “When Can in-Context Learning Generalize Out of Task Distribution?”
Guan, Wu, Zhao, et al. 2025. “Attention Mechanisms Perspective: Exploring LLM Processing of Graph-Structured Data.”
Hasson, Nastase, and Goldstein. 2020. “Direct Fit to Nature: An Evolutionary Perspective on Biological and Artificial Neural Networks.” Neuron.
Jeffares, Curth, and van der Schaar. 2024. “Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & Beyond.”
Khatri, Laakkonen, Liu, et al. 2024. “On the Anatomy of Attention.”
Kleinberg, and Mullainathan. 2024. “Language Generation in the Limit.”
Madabushi, Torgbi, and Bonial. 2025. “Neither Stochastic Parroting nor AGI: LLMs Solve Tasks Through Context-Directed Extrapolation from Training Data Priors.”
Misra, and Mahowald. 2024. “Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs.”
Morris, Sitawarin, Guo, et al. 2025. “How Much Do Language Models Memorize?”
Power, Burda, Edwards, et al. 2022. “Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.”
Su, Tsai, Wang, et al. 2009. “Subgroup Analysis via Recursive Partitioning.” In Journal of Machine Learning Research.
Vaswani, Shazeer, Parmar, et al. 2017. “Attention Is All You Need.” arXiv:1706.03762 [Cs].
Vasylenko, Treviso, and Martins. 2025. “Long-Context Generalization with Sparse Attention.”
Veličković, Cucurull, Casanova, et al. 2018. “Graph Attention Networks.”
Veličković, Perivolaropoulos, Barbero, et al. 2025. “Softmax Is Not Enough (for Sharp Size Generalisation).” In.
Webb, Dulberg, Frankland, et al. 2020. “Learning Representations That Support Extrapolation.” In Proceedings of the 37th International Conference on Machine Learning.
Wilson, and Izmailov. 2020. “Bayesian Deep Learning and a Probabilistic Perspective of Generalization.”
Xu, Zhang, Li, et al. 2021. “How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks.”
Yadlowsky, Doshi, and Tripuraneni. 2023. “Can Transformer Models Generalize Via In-Context Learning Beyond Pretraining Data?” In.
Zhang, Bengio, Hardt, et al. 2021. “Understanding Deep Learning (Still) Requires Rethinking Generalization.” Communications of the ACM.