Scaled Dot-Product Attention & Transformer Architectures
Mathematical mechanics of Scaled Dot-Product Attention, Multi-Head Attention, Positional Encodings, and Transformer encoder-decoder topologies.
01 — Notebook Information & Scope
“Attention is All You Need” (Vaswani et al., 2017) discarded sequential recurrent state propagation (RNNs/LSTMs) in favor of parallelized matrix dot-product attention, allowing models to route representations directly between arbitrary token pairs regardless of distance.
- Domain: Artificial Intelligence & Machine Learning
- Subject: Deep Learning & Transformers
- Core Reference Paper: Vaswani et al. (2017), “Attention is All You Need”
02 — Scaled Dot-Product Attention
Given an input matrix of Query vectors $Q$, Key vectors $K$, and Value vectors $V$ of dimension $d_k$:
Scaled Dot-Product Attention Equation
FORMULAThe inner product Q K^T measures semantic affinity between token pairs. Softmax normalizes affinities into a probability distribution over keys, weighting the linear combination of values V.
$Q \in \mathbb{R}^{n \times d_k}$Matrix of query vectors representing tokens seeking context$K \in \mathbb{R}^{m \times d_k}$Matrix of key vectors representing tokens offering context$V \in \mathbb{R}^{m \times d_v}$Matrix of value vectors containing the token representation payloads$\sqrt{d_k}$Scaling factor to prevent dot products from growing excessively large for large dimensions, which would push softmax into near-zero gradient regions03 — Multi-Head Attention (MHA)
Rather than performing a single attention function with $d_{model}$-dimensional queries, keys, and values, Multi-Head Attention projects $Q, K, V$ linearly $h$ times with different learned parameter matrices:
Multi-Head Attention (MHA) Formulation
FORMULAProjects queries, keys, and values h times with distinct parameter projections, enabling the model to attend to multiple representation subspaces simultaneously.
- Allows the model to jointly attend to information from different representation subspaces (e.g., one head tracking syntactic subject-verb dependencies, another tracking coreference resolution).