🌏 中文版
Version note: The slides are CS231N Spring 2026 lecture_3.pdf (121 pages, footer dated April 7, 2026). The recording is the Spring 2025 Lecture 3, because 2026 recordings are on Canvas for enrolled students only. The two may differ. The 2025 deck has 119 pages, and the keywords I spot-checked (AdaGrad, AdamW, L-BFGS, warmup) appear in both, but I did not compare them page by page. All facts were checked against official materials on 2026-09-30. Access level A3 (defined in the Global AI/CS Course Map).
Series: Previous L2: Image Classification, kNN, and Linear Classifiers | Next L4: Neural Networks and Backpropagation | Series overview
Lecture 2 of CS231N handed us three things: a dataset of (x, y) pairs, a score function f(x, W) = Wx + b, and a softmax loss that says how good the scores are. Lecture 3 asks the next question: given a loss, how do we find a W that makes it small?
The 2026 schedule lists four topics for this lecture: regularization, stochastic gradient descent, Momentum/AdaGrad/Adam, and learning rate schedules. It sits in the "Deep Learning Basics" unit. Every model later in the course (CNNs, RNNs, Transformers, diffusion) reuses this toolbox.
This is the first hard post in the series, so it follows five layers: the setting, the intuition, the mechanics (formulas in collapsible blocks), how it connects to the course's models, and where to go deeper.
The setting: walking downhill without a map
At the start of the optimization section, the slides show a photo of a valley and a walking figure. One way to read it: the loss is altitude, W is where you stand, and the goal is the valley floor. With tens of thousands of parameters you can't see the whole terrain. You can only compute the slope under your feet.
The first strategy is deliberately bad: try random Ws and keep the one with the lowest loss. The slides report 15.5% test accuracy this way, with a note that the state of the art is about 99.7%. The optimization-1 course notes use the same example and add that CIFAR-10 has ten classes, so random guessing gets 10%, and 15.5% isn't that bad.
The second strategy is to follow the slope. That's the thread of the whole lecture.
Before it gets there, L3 deals with a different question: there's more than one valley floor, so which one do you want?
Intuition 1: regularization writes your preferences into the loss
The full objective on the slides is:
$$ L(W) = \frac{1}{N}\sum_{i=1}^{N} L_i(f(x_i, W), y_i) + \lambda R(W) $$
The first term is the data loss, which asks the predictions to match the training data. The second is regularization, which keeps the model from doing too well on the training data. λ is the regularization strength, a hyperparameter.
Why stop a model from doing well? The slides use a toy example: a wiggly curve f1 that passes through every training point, and a smooth line f2. When new data arrives, f2 usually does better. The slide cites Occam's razor: among competing hypotheses that explain the data, prefer the simplest.
The slides give three reasons to regularize:
- Express preferences over weights. L2 regularization, for example, likes to "spread out" the weights instead of concentrating them in a few dimensions.
- Keep the model simple so it works on test data.
- Improve optimization by adding curvature.
The common forms come in two groups. The simple ones are L2, L1, and elastic net (L1 + L2). The more complex ones are dropout, batch normalization, stochastic depth, and fractional pooling. This lecture only names the second group. You implement dropout and batch norm in A2.
The slides leave you a question. With input x = [1, 1, 1, 1], two weight vectors, w1 = [1, 0, 0, 0] and w2 = [0.25, 0.25, 0.25, 0.25], both give a score of 1. Which one does L2 regularization, R(W) = ΣΣ W², prefer? The slide boxes w2 and says L2 likes to spread out the weights. Which one L1 prefers is left as an open question on the slide.
Intuition 2: the gradient tells you which way to walk
In one dimension the derivative is a slope. In many dimensions the gradient is the vector of partial derivatives. Three lines from the slides are worth keeping: the slope in any direction is the dot product of that direction with the gradient, and the direction of steepest descent is the negative gradient.
There are two ways to compute it:
| Method | How | The slides' verdict |
|---|---|---|
| Numerical gradient | Nudge each dimension by a small h and see how the loss changes | Approximate, slow, easy to write |
| Analytic gradient | Use calculus to write ∇W L directly | Exact, fast, error-prone |
In practice you use both: always train with the analytic gradient, and check your implementation with the numerical one. This is called a gradient check. Every place in A1 where you write a gradient comes with a gradient-check cell.
Once you have a gradient, gradient descent repeatedly takes a small step in the negative gradient direction. When N is large, computing the full loss for every step is too expensive, so you estimate the gradient from a small batch. The slides list 32, 64, and 128 as common sizes. That's SGD.
Intuition 3: three problems with SGD, and the fixes
The slides sort SGD's failure cases into three problems.
Problem 1: steep in one direction, shallow in another. SGD jitters back and forth along the steep direction and makes slow progress along the shallow one. The slide notes that such a loss has a high condition number, meaning the ratio of the Hessian's largest to smallest singular value is large.
Problem 2: local minima and saddle points. The gradient is zero, so gradient descent gets stuck. Citing Dauphin et al. (NIPS 2014), the slides point out that saddle points are much more common than local minima in high dimensions.
Problem 3: noisy gradients. The gradient comes from a minibatch, so it's an estimate.
Each optimizer that follows responds to some of these:
| Optimizer | What it does | Which problem |
|---|---|---|
| SGD + Momentum | Accumulates past gradients into a "velocity" and keeps moving in the general direction | Jitter, saddle points, noise |
| RMSProp | Scales each dimension's step by its history of squared gradients | Damps steep directions, speeds up flat ones |
| Adam | Momentum + RMSProp, plus bias correction | Both |
| AdamW | Adam, but weight decay stays out of the moment estimates | How regularization interacts with the optimizer |
AdaGrad is on the schedule, but in the 2026 slides its full treatment sits in "Appendix / Enrichment Material (Slides from Previous Years)". The main deck only labels the second-moment line on the Adam slide as "AdaGrad / RMSProp". The key question in the appendix is what happens to AdaGrad's step size over a long time. The answer: it decays to zero, because AdaGrad sums every squared gradient it has ever seen. That's why RMSProp is called "leaky AdaGrad": it lets the old sum decay.
Mechanics: what the updates look like
The code and numbers below are copied from the slides.
SGD and SGD + Momentum
SGD:
$$x_{t+1} = x_t - \alpha \nabla f(x_t)$$
SGD + Momentum:
$$v_{t+1} = \rho v_t + \nabla f(x_t), \qquad x_{t+1} = x_t - \alpha v_{t+1}$$
From the slides: the velocity is a running mean of gradients, ρ gives "friction", and typical values are 0.9 or 0.99. The source is Sutskever et al. (ICML 2013). The slides also warn that you'll see other formulations, but they're equivalent and produce the same sequence of x.
RMSProp
grad_squared = 0
while True:
dx = compute_gradient(x)
grad_squared = decay_rate * grad_squared + (1 - decay_rate) * dx * dx
x -= learning_rate * dx / (np.sqrt(grad_squared) + 1e-7)
Attributed to Tieleman and Hinton, 2012. Steep directions (large grad_squared) get smaller steps and flat directions get relatively larger ones. This is what "per-parameter learning rates" means.
Adam (full form)
first_moment = 0
second_moment = 0
for t in range(1, num_iterations):
dx = compute_gradient(x)
first_moment = beta1 * first_moment + (1 - beta1) * dx # Momentum
second_moment = beta2 * second_moment + (1 - beta2) * dx * dx # AdaGrad / RMSProp
first_unbias = first_moment / (1 - beta1 ** t) # Bias correction
second_unbias = second_moment / (1 - beta2 ** t)
x -= learning_rate * first_unbias / (np.sqrt(second_unbias) + 1e-7)
Why bias correction? The slides first show an "almost Adam" version and ask what happens at the first timestep. Both moments start at zero, so the second moment is tiny for the first few steps. Dividing by it makes the first step huge. Bias correction compensates for estimates that start at zero. The source is Kingma and Ba (ICLR 2015).
The slides' starting values: beta1 = 0.9, beta2 = 0.999, learning_rate = 1e-3 or 5e-4, "a great starting point for many models".
AdamW: where weight decay goes
The slides ask how regularization (e.g., L2) interacts with the optimizer. The answer: it depends.
- Standard Adam: the L2 term is part of the gradient dx, so it enters the first- and second-moment estimates and gets rescaled by the RMSProp term.
- AdamW: the weight decay term is added after the moments are computed, directly on the final update line.
The slide includes a plot of ImageNet accuracy vs. training epoch from a 2018 fast.ai post.
Learning rate schedules and warmup
SGD, Momentum, RMSProp, Adam, and AdamW all have a learning rate. The slides show several learning rate curves and ask which one is best. The answer: in reality, any of them could be a good learning rate. What differs is when you use it.
| Schedule | Formula | Example on the slides |
|---|---|---|
| Step | Drop at a few fixed points | ResNets multiply the LR by 0.1 after epochs 30, 60, and 90 |
| Cosine | $\alpha_t = \frac{1}{2}\alpha_0(1+\cos(t\pi/T))$ | Cites SGDR, GPT, SlowFast, Sparse Transformers |
| Linear | $\alpha_t = \alpha_0(1 - t/T)$ | Cites BERT |
| Inverse sqrt | $\alpha_t = \alpha_0/\sqrt{t}$ | Cites Attention Is All You Need |
α₀ is the initial learning rate, α_t the rate at epoch t, and T the total number of epochs.
Linear warmup: a high initial learning rate can make the loss explode, so increase it linearly from 0 over roughly the first 5,000 iterations. The same slide gives a rule of thumb: if you increase the batch size by N, scale the initial learning rate by N too, citing Goyal et al. (2017).
Why second-order optimization doesn't fit deep learning
First-order methods build a linear approximation from the gradient and step toward its minimum. Second-order methods add the Hessian to build a quadratic approximation and jump straight to its minimum (Newton's method).
The slides' answer is blunt: the Hessian has O(N²) entries, inverting it takes O(N³), and N is tens or hundreds of millions of parameters. The appendix adds BFGS and L-BFGS. L-BFGS usually works very well in full-batch, deterministic settings, but it doesn't transfer well to minibatches. The slides call adapting second-order methods to large-scale stochastic settings an active research area.
Back to the models: how the course uses this
The slides' "In practice" list has three items:
- Adam(W) is a good default choice in many cases, and it often works fine even with a constant learning rate.
- SGD + Momentum can outperform Adam, but it may need more tuning of the learning rate and schedule.
- If you can afford full-batch updates, look beyond first-order methods.
These carry through to the final project. Q5 of A1 has you implement sgd_momentum, rmsprop, and adam in cs231n/optim.py and use them to train multi-layer fully connected networks. The starter code's Adam defaults match the slides exactly: beta1 = 0.9, beta2 = 0.999, learning_rate = 1e-3. The same notebook has an inline question on why AdaGrad's updates keep shrinking and whether Adam has the same issue. That's the appendix material.
The end of L3 already sets up the next lecture. A linear classifier can't separate data where red points sit inside a ring of blue points. Switch to polar coordinates and a line separates them. Rather than hand-design such transforms, let the model learn them: that's a two-layer neural network. But once the network gets deeper, how do you compute the gradient? That's L4.
Slides vs. course notes
The schedule links this lecture to the optimization-1 notes. They cover a different range than the slides:
- The notes use the multiclass SVM loss; the slides use softmax. The notes visualize the SVM loss landscape and point out that it's convex, then warn that with neural networks the objective becomes non-convex, a bumpy terrain.
- Between random search and following the gradient, the notes add a "random local search" strategy.
- The notes stop at minibatch gradient descent. They have no Momentum, Adam, or learning rate schedules. Those live in the "Parameter updates" section of another set of notes, neural-networks-3, which is also where the Nesterov slides in the appendix point.
Going deeper
- Intuition for momentum: the schedule lists Distill's Why Momentum Really Works as a suggested reading for L4, but it's most relevant here. You can drag ρ and the learning rate and watch the oscillation change.
- Fuller notes on parameter updates: neural-networks-3 covers Nesterov momentum, learning rate annealing, AdaGrad/RMSProp, hyperparameter search, and practical gradient checking.
- The same topic from another angle: CMU 11-785 Lecture 8: Optimizers and Regularization and the MIT 6.7960 guide.
- Shaky on vector calculus or probability: Stanford CS229 guide, Stanford CS109 guide.
How to self-study it
- Watch the 2025 Lecture 3 recording with the 2026 slides open. Where they disagree, go with the slides.
- Read the numerical gradient and gradient check sections of optimization-1. That's where A1 most often goes wrong.
- Copy the RMSProp and Adam code from the slides into a notebook, run them on a narrow 2-D bowl, and plot the trajectories.
One thing you can do tonight: in numpy, define f(x, y) = x² + 20y², run SGD and SGD + Momentum (ρ = 0.9) from the same starting point for 50 steps, and plot both paths. You'll see what "jitter along the steep direction, crawl along the shallow one" actually looks like.
References
- CS231N course home (Spring 2026) — instructors, grading, Canvas-only recordings
- CS231N Spring 2026 schedule — L3 topics, optimization-1 link, L4 suggested readings
- Lecture 3 slides: Regularization and Optimization (2026) — source of every formula, code snippet, and number in this post
- Lecture 3 slides (2025) — same year as the recording, for comparison
- Stanford CS231N Spring 2025 Lecture 3 recording — public Stanford Online recording
- CS231N notes: Optimization: Stochastic Gradient Descent — SVM loss landscape, three search strategies, minibatch GD
- CS231N notes: Neural Networks Part 3 — the Parameter updates section
- Assignment 1 (2026) — Q5 implements SGD+Momentum, RMSProp, Adam
- Kingma & Ba, Adam: A Method for Stochastic Optimization
- Sutskever et al., On the importance of initialization and momentum in deep learning (ICML 2013)
- Dauphin et al., Identifying and attacking the saddle point problem (NIPS 2014)
- Goyal et al., Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- fast.ai, AdamW and Super-convergence is now the fastest way to train neural nets
- Goh, Why Momentum Really Works (Distill, 2017)
Loading...