CANONICAL EDITION
VERIFIED 100%ACADEMIC MASTER
Artificial Intelligence & Systems•Transformers & Foundation Models

Scaled Dot-Product Attention & Transformer Architectures

Mathematical mechanics of Scaled Dot-Product Attention, Multi-Head Attention, Positional Encodings, and Transformer encoder-decoder topologies.

Faculty Reference: Personal Notes
Updated: 2026-10-05
Format: Canonical Markdown/MDX

01 — Notebook Information & Scope

The Attention Revolution

“Attention is All You Need” (Vaswani et al., 2017) discarded sequential recurrent state propagation (RNNs/LSTMs) in favor of parallelized matrix dot-product attention, allowing models to route representations directly between arbitrary token pairs regardless of distance.

  • Domain: Artificial Intelligence & Machine Learning
  • Subject: Deep Learning & Transformers
  • Core Reference Paper: Vaswani et al. (2017), “Attention is All You Need”

02 — Scaled Dot-Product Attention

Given an input matrix of Query vectors $Q$, Key vectors $K$, and Value vectors $V$ of dimension $d_k$:

Scaled Dot-Product Attention Equation

FORMULA

The inner product Q K^T measures semantic affinity between token pairs. Softmax normalizes affinities into a probability distribution over keys, weighting the linear combination of values V.

$$\mathrm{Attention}(Q, K, V) = \mathrm{softmax}\left(\frac{Q K^T}{\sqrt{d_k}}\right) V$$
Variable Definitions & Units:
$Q \in \mathbb{R}^{n \times d_k}$Matrix of query vectors representing tokens seeking context
$K \in \mathbb{R}^{m \times d_k}$Matrix of key vectors representing tokens offering context
$V \in \mathbb{R}^{m \times d_v}$Matrix of value vectors containing the token representation payloads
$\sqrt{d_k}$Scaling factor to prevent dot products from growing excessively large for large dimensions, which would push softmax into near-zero gradient regions

03 — Multi-Head Attention (MHA)

Rather than performing a single attention function with $d_{model}$-dimensional queries, keys, and values, Multi-Head Attention projects $Q, K, V$ linearly $h$ times with different learned parameter matrices:

Multi-Head Attention (MHA) Formulation

FORMULA

Projects queries, keys, and values h times with distinct parameter projections, enabling the model to attend to multiple representation subspaces simultaneously.

$$\mathrm{MultiHead}(Q, K, V) = \mathrm{Concat}(\mathrm{head}_1, \dots, \mathrm{head}_h) W^O, \quad \text{where } \mathrm{head}_i = \mathrm{Attention}(Q W_i^Q, K W_i^K, V W_i^V)$$
  • Allows the model to jointly attend to information from different representation subspaces (e.g., one head tracking syntactic subject-verb dependencies, another tracking coreference resolution).

04 — Active Recall Flashcards & Conceptual Quiz

🗂 Flashcard • Key ConceptClick to Flip
Why is the dot-product Q * K^T divided by sqrt(d_k) in Scaled Dot-Product Attention?
Reveal Definition / Answer ↓
For large values of d_k, the dot products grow large in magnitude, pushing the softmax function into regions with extremely small gradients (vanishing gradient problem). Scaling by 1 / sqrt(d_k) restores unit variance.
Conceptual Check / QuizActive Recall

Why are Positional Encodings required in Transformer architectures?