🌏 中文版
Version note: This post mainly follows the Lecture 8 slides linked from the Spring 2026 CS231N schedule (124 pages, downloaded and checked on 2026-09-30), plus the RNNs & Transformers review slides from the 5/1 section, whose cover says they were copied from the 2025 version. For video, watch Spring 2025's Lecture 8: Attention and Transformers; 2026 recordings are on Canvas for enrolled students only. The two years' slides are mostly the same, but the 2026 deck adds a page each on RoPE and QK-Norm, so the video won't cover those two. Access level A3.
Series: previous A2 guide: BatchNorm, Dropout, CNNs, PyTorch, and RNN Captioning | next L9: Object Detection, Image Segmentation, and Visualization | Series overview
Through Lecture 7, CS231N has two kinds of structure: convolution for grids and RNNs for sequences. Lecture 8 introduces a third, and it went on to take over the territory of the other two. The summary slide says it outright: Transformers are the backbone of all large AI models today, used for language, vision, speech, and more.
The lecture follows one line: where attention came from → abstracting it into a general operation → building the Transformer out of it → turning images into something a Transformer can consume. This post follows the same line.
The problem: an RNN translator squeezed through one vector
The slides open with seq2seq translation, turning "we see the sky" into the Italian "vediamo il cielo". The encoder RNN reads the whole English sentence and hands the decoder only its final hidden state, a summary c of the sentence. The longer the sentence, the more has to fit into that one fixed-size vector.
Bahdanau et al. 2015 let the decoder look back at every encoder hidden state each time it emits a word:
- Compute an alignment score e_{t,i} between the previous decoder state s_{t−1} and each encoder state h_i. The slides say f_att is a linear layer.
- Softmax turns the scores into weights a_{t,i}.
- Sum the h_i with those weights to get a context vector c_t for this step.
Each timestep uses a different context vector, and you can plot the weights to see which source words the model attends to while translating.
The slides then point out a general operation hiding here. Decoder states act as queries, encoder states act as data vectors, and each query looks at all the data vectors and produces one output. None of this needs an RNN.
Intuition: attention is an operation on sets
With the RNN removed, an attention layer takes a set of query vectors Q and a set of data vectors X. The slides build up the full shapes page by page:
- Project keys from X: K = XW_K; project values: V = XW_V
- Similarities E = QKᵀ / √D_Q, each entry a dot product between one query and one key
- Weights A = softmax(E), giving each query a distribution over the keys
- Output Y = AV, each output a weighted sum of values
When queries and data come from different sources, this is cross-attention. When the queries are computed from the same inputs (Q = XW_Q), it's self-attention: each input produces one output, and that output mixes information from all inputs. In practice the Q, K, and V projections are often fused into one matmul: [Q K V] = X[W_Q W_K W_V].
Full self-attention shapes (slide 47)
input X [N × D_in]
Q = X W_Q [N × D_out]
K = X W_K [N × D_out]
V = X W_V [N × D_out]
E = Q Kᵀ / √D_Q [N × N]
A = softmax(E) [N × N] each query normalized over all keys
Y = A V [N × D_out] Y_i = Σ_j A_ij V_j
The slides note that almost always D_Q = D_V = D_out.
Mechanism: four properties that shape the Transformer
1. It doesn't know order
Permute the inputs, and Q, K, V, the similarities, the weights, and the outputs are all permuted the same way. Nothing else changes. The slides write this as F(σ(X)) = σ(F(X)) and call it permutation equivariance.
That's both a strength and a problem. Self-attention is a natural fit for sets, but on a sentence it can't tell "dog bites man" from "man bites dog". The slides give two fixes:
- Add a positional encoding to each input, a fixed function of the position index.
- RoPE (Su et al. 2021): map positions to angles and rotate queries and keys, so their dot product depends only on relative position. This page is new in 2026; the 2025 deck doesn't have it.
2. Masks control what it can see
Masked self-attention overrides the similarities it shouldn't see with −∞, so their softmax weights become 0. Language models use this so each token sees only the tokens before it and can't peek at the answer.
3. You can run several copies in parallel
Multi-head self-attention runs H copies of self-attention in parallel, one per head, then concatenates their outputs and projects back to the original dimension.
4. The whole layer is four matrix multiplies
The slides break multi-head self-attention into four steps:
- QKV projection: [N × D] times [D × 3HD_H]
- QK similarity: giving [H × N × N]
- Weighting V: giving [H × N × D_H]
- Output projection: [N × HD_H] times [HD_H × D]
The trouble is step 2's H × N × N attention matrix. The slides' example: with N = 100K and H = 64, that matrix alone takes 1.192 TB, more than a GPU holds. The fix is Flash Attention, which computes steps 2 and 3 together without storing the full attention matrix, making large N feasible.
Three ways to process sequences
The slides compare RNNs, convolution, and self-attention side by side. This table is the page most worth remembering from the lecture:
| RNN | Convolution | Self-attention | |
|---|---|---|---|
| Works on | 1D ordered sequences | N-dimensional grids | Sets of vectors |
| Long sequences | Good in theory; O(N) compute and memory | Bad; needs many stacked layers to see far | Good; each output depends directly on all inputs |
| Parallelism | No; hidden states computed sequentially | Yes | Yes; it's just four matmuls |
| Cost | — | — | Expensive: O(N²) compute, O(N) memory |
The 5/1 section slides add another angle. RNNs have a strong inductive bias with temporal structure built in; Transformers have a weak one and must learn it from data.
Back to the model: the Transformer block
A block of the Transformer (Vaswani et al. 2017), from bottom to top:
- Multi-head self-attention: where all vectors interact
- Residual connection + LayerNorm: LayerNorm normalizes each vector on its own
- MLP: usually two layers, classically D → 4D → D, also called an FFN, applied to each vector independently
- Another residual connection + LayerNorm
A Transformer is just a stack of identical blocks. The slides stress three points. Self-attention is the only place vectors interact. LayerNorm and the MLP work on each vector independently. Most of the compute is just 6 matmuls, 4 in self-attention and 2 in the MLP, which makes the architecture highly scalable and parallelizable. The slides add that it hasn't changed much since 2017; it has mostly gotten a lot bigger.
For language: LLMs
At the input, learn a [V × D] embedding matrix that turns words into vectors. Inside each block, use masked attention so each token sees only earlier tokens. At the output, learn a [D × V] projection matrix that turns each D-dimensional vector into scores over the vocabulary.
For images: ViT
ViT (Dosovitskiy et al., ICLR 2021, titled "An Image is Worth 16x16 Words") answers the question: an image isn't a sequence of words, so how does it become Transformer input? The slides' steps:
- Take an input image, e.g. 224×224×3.
- Break it into patches, e.g. 16×16×3.
- Flatten each patch (16×16×3 = 768 values) and apply a linear projection to D dimensions.
- Add positional encoding so the Transformer knows each patch's 2D position.
- Use no masking: every patch can look at every other patch.
- The Transformer outputs one vector per patch; average-pool the N vectors into one, then apply a linear layer D → C to predict class scores.
At step 3 the slides pause to ask: is there another way to describe this operation? The answer is a 16×16 convolution with stride 16, 3 input channels, and D output channels. That one line connects ViT to the CNNs of earlier lectures: ViT's first layer is really a large-stride convolution, and self-attention takes over from there.
Four common changes since 2017
The last section lists adjustments that have become standard in modern Transformers:
- Pre-Norm: the original puts LayerNorm outside the residual connection, which the slides call "kind of weird" because the model can't learn the identity function. The fix moves normalization inside the residual branch.
- QK-Norm: normalize queries and keys before computing similarities, which prevents gradient spikes and stabilizes training. This page is also new in 2026.
- SwiGLU: replace the classic MLP with Y = (σ(XW₁) ⊙ XW₂)W₃. Setting the hidden size to H = 8D/3 keeps the parameter count the same (Shazeer 2020).
- Mixture of Experts (MoE): learn E sets of MLP weights per block and route each token to only A of them. Parameters grow by a factor of E while compute grows only with A (Shazeer et al. 2017).
The 2025 deck had RMSNorm in this section. The 2026 Lecture 8 drops that page and moves it to the Transformer recap at the start of Lecture 9.
Going deeper
- Suggested readings on the schedule: the original Attention Is All You Need paper, Lilian Weng's Attention? Attention!, Jay Alammar's The Illustrated Transformer, and the ViT paper.
- Hands-on: the section 5 slides end with a Colab notebook. The first question in this series' A3 guide swaps A2's RNN for a Transformer in image captioning.
- Self-check: without the slides, write out the matrix shape at every step of self-attention, and explain why ViT's patch embedding is a convolution.
Further reading
These courses cover the same architecture from the language model side. This post's vision thread doesn't depend on them:
- CS224N: Transformers
- CME295: Transformers
- Building a Transformer language model from scratch: CS336 guide
References
- CS231N Lecture 8 slides (Spring 2026) — source of every figure, shape, and example in this post
- CS231N Section 5: RNNs & Transformers slides — RNN versus Transformer comparison
- CS231N schedule (Spring 2026) — the 4/23 lecture and suggested readings
- CS231N Lecture 8 slides (Spring 2025) — for comparison with 2026
- Spring 2025 Lecture 8 recording
- Vaswani et al., Attention Is All You Need (NeurIPS 2017)
- Dosovitskiy et al., An Image is Worth 16x16 Words (ICLR 2021)
- Bahdanau, Cho & Bengio, Neural Machine Translation by Jointly Learning to Align and Translate (ICLR 2015)
- Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding (2021)
- Dao et al., FlashAttention (2022)
- Shazeer, GLU Variants Improve Transformer (2020)
- Shazeer et al., Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (2017)
- Lilian Weng, Attention? Attention!
- Jay Alammar, The Illustrated Transformer
Loading...