Skip to content

CS231N L8: Attention, Transformers, and ViT

Sep 30, 20261 min
TL;DRLecture 8 of CS231N Spring 2026 starts from the bottleneck in RNN translation models, abstracts attention into an operation on sets of vectors, builds up to self-attention, masking, and multiple heads, and shows the whole layer is four matrix multiplies. A Transformer block is self-attention, LayerNorm, residual connections, and an MLP; ViT turns a 224×224 image into 16×16 patches used as tokens. The lecture closes with four common post-2017 changes: Pre-Norm, QK-Norm, SwiGLU, and MoE.

🌏 中文版

Version note: This post mainly follows the Lecture 8 slides linked from the Spring 2026 CS231N schedule (124 pages, downloaded and checked on 2026-09-30), plus the RNNs & Transformers review slides from the 5/1 section, whose cover says they were copied from the 2025 version. For video, watch Spring 2025's Lecture 8: Attention and Transformers; 2026 recordings are on Canvas for enrolled students only. The two years' slides are mostly the same, but the 2026 deck adds a page each on RoPE and QK-Norm, so the video won't cover those two. Access level A3.

Series: previous A2 guide: BatchNorm, Dropout, CNNs, PyTorch, and RNN Captioning | next L9: Object Detection, Image Segmentation, and Visualization | Series overview

Through Lecture 7, CS231N has two kinds of structure: convolution for grids and RNNs for sequences. Lecture 8 introduces a third, and it went on to take over the territory of the other two. The summary slide says it outright: Transformers are the backbone of all large AI models today, used for language, vision, speech, and more.

The lecture follows one line: where attention came from → abstracting it into a general operation → building the Transformer out of it → turning images into something a Transformer can consume. This post follows the same line.

The problem: an RNN translator squeezed through one vector

The slides open with seq2seq translation, turning "we see the sky" into the Italian "vediamo il cielo". The encoder RNN reads the whole English sentence and hands the decoder only its final hidden state, a summary c of the sentence. The longer the sentence, the more has to fit into that one fixed-size vector.

Bahdanau et al. 2015 let the decoder look back at every encoder hidden state each time it emits a word:

  1. Compute an alignment score e_{t,i} between the previous decoder state s_{t−1} and each encoder state h_i. The slides say f_att is a linear layer.
  2. Softmax turns the scores into weights a_{t,i}.
  3. Sum the h_i with those weights to get a context vector c_t for this step.

Each timestep uses a different context vector, and you can plot the weights to see which source words the model attends to while translating.

The slides then point out a general operation hiding here. Decoder states act as queries, encoder states act as data vectors, and each query looks at all the data vectors and produces one output. None of this needs an RNN.

Intuition: attention is an operation on sets

With the RNN removed, an attention layer takes a set of query vectors Q and a set of data vectors X. The slides build up the full shapes page by page:

  • Project keys from X: K = XW_K; project values: V = XW_V
  • Similarities E = QKᵀ / √D_Q, each entry a dot product between one query and one key
  • Weights A = softmax(E), giving each query a distribution over the keys
  • Output Y = AV, each output a weighted sum of values

When queries and data come from different sources, this is cross-attention. When the queries are computed from the same inputs (Q = XW_Q), it's self-attention: each input produces one output, and that output mixes information from all inputs. In practice the Q, K, and V projections are often fused into one matmul: [Q K V] = X[W_Q W_K W_V].

Full self-attention shapes (slide 47)
input X          [N × D_in]
Q = X W_Q        [N × D_out]
K = X W_K        [N × D_out]
V = X W_V        [N × D_out]
E = Q Kᵀ / √D_Q   [N × N]
A = softmax(E)   [N × N]      each query normalized over all keys
Y = A V          [N × D_out]   Y_i = Σ_j A_ij V_j

The slides note that almost always D_Q = D_V = D_out.

Mechanism: four properties that shape the Transformer

1. It doesn't know order

Permute the inputs, and Q, K, V, the similarities, the weights, and the outputs are all permuted the same way. Nothing else changes. The slides write this as F(σ(X)) = σ(F(X)) and call it permutation equivariance.

That's both a strength and a problem. Self-attention is a natural fit for sets, but on a sentence it can't tell "dog bites man" from "man bites dog". The slides give two fixes:

  • Add a positional encoding to each input, a fixed function of the position index.
  • RoPE (Su et al. 2021): map positions to angles and rotate queries and keys, so their dot product depends only on relative position. This page is new in 2026; the 2025 deck doesn't have it.

2. Masks control what it can see

Masked self-attention overrides the similarities it shouldn't see with −∞, so their softmax weights become 0. Language models use this so each token sees only the tokens before it and can't peek at the answer.

3. You can run several copies in parallel

Multi-head self-attention runs H copies of self-attention in parallel, one per head, then concatenates their outputs and projects back to the original dimension.

4. The whole layer is four matrix multiplies

The slides break multi-head self-attention into four steps:

  1. QKV projection: [N × D] times [D × 3HD_H]
  2. QK similarity: giving [H × N × N]
  3. Weighting V: giving [H × N × D_H]
  4. Output projection: [N × HD_H] times [HD_H × D]

The trouble is step 2's H × N × N attention matrix. The slides' example: with N = 100K and H = 64, that matrix alone takes 1.192 TB, more than a GPU holds. The fix is Flash Attention, which computes steps 2 and 3 together without storing the full attention matrix, making large N feasible.

Three ways to process sequences

The slides compare RNNs, convolution, and self-attention side by side. This table is the page most worth remembering from the lecture:

RNNConvolutionSelf-attention
Works on1D ordered sequencesN-dimensional gridsSets of vectors
Long sequencesGood in theory; O(N) compute and memoryBad; needs many stacked layers to see farGood; each output depends directly on all inputs
ParallelismNo; hidden states computed sequentiallyYesYes; it's just four matmuls
Cost——Expensive: O(N²) compute, O(N) memory

The 5/1 section slides add another angle. RNNs have a strong inductive bias with temporal structure built in; Transformers have a weak one and must learn it from data.

Back to the model: the Transformer block

A block of the Transformer (Vaswani et al. 2017), from bottom to top:

  1. Multi-head self-attention: where all vectors interact
  2. Residual connection + LayerNorm: LayerNorm normalizes each vector on its own
  3. MLP: usually two layers, classically D → 4D → D, also called an FFN, applied to each vector independently
  4. Another residual connection + LayerNorm

A Transformer is just a stack of identical blocks. The slides stress three points. Self-attention is the only place vectors interact. LayerNorm and the MLP work on each vector independently. Most of the compute is just 6 matmuls, 4 in self-attention and 2 in the MLP, which makes the architecture highly scalable and parallelizable. The slides add that it hasn't changed much since 2017; it has mostly gotten a lot bigger.

For language: LLMs

At the input, learn a [V × D] embedding matrix that turns words into vectors. Inside each block, use masked attention so each token sees only earlier tokens. At the output, learn a [D × V] projection matrix that turns each D-dimensional vector into scores over the vocabulary.

For images: ViT

ViT (Dosovitskiy et al., ICLR 2021, titled "An Image is Worth 16x16 Words") answers the question: an image isn't a sequence of words, so how does it become Transformer input? The slides' steps:

  1. Take an input image, e.g. 224×224×3.
  2. Break it into patches, e.g. 16×16×3.
  3. Flatten each patch (16×16×3 = 768 values) and apply a linear projection to D dimensions.
  4. Add positional encoding so the Transformer knows each patch's 2D position.
  5. Use no masking: every patch can look at every other patch.
  6. The Transformer outputs one vector per patch; average-pool the N vectors into one, then apply a linear layer D → C to predict class scores.

At step 3 the slides pause to ask: is there another way to describe this operation? The answer is a 16×16 convolution with stride 16, 3 input channels, and D output channels. That one line connects ViT to the CNNs of earlier lectures: ViT's first layer is really a large-stride convolution, and self-attention takes over from there.

Four common changes since 2017

The last section lists adjustments that have become standard in modern Transformers:

  • Pre-Norm: the original puts LayerNorm outside the residual connection, which the slides call "kind of weird" because the model can't learn the identity function. The fix moves normalization inside the residual branch.
  • QK-Norm: normalize queries and keys before computing similarities, which prevents gradient spikes and stabilizes training. This page is also new in 2026.
  • SwiGLU: replace the classic MLP with Y = (σ(XW₁) ⊙ XW₂)W₃. Setting the hidden size to H = 8D/3 keeps the parameter count the same (Shazeer 2020).
  • Mixture of Experts (MoE): learn E sets of MLP weights per block and route each token to only A of them. Parameters grow by a factor of E while compute grows only with A (Shazeer et al. 2017).

The 2025 deck had RMSNorm in this section. The 2026 Lecture 8 drops that page and moves it to the Transformer recap at the start of Lecture 9.

Going deeper

Further reading

These courses cover the same architecture from the language model side. This post's vision thread doesn't depend on them:

References