L2 gave us a score function and a loss. L3 answers how to find a good W. The first half covers regularization: add λR(W) next to the data loss so the model does not fit the training data too well. The second half is a lineage of optimizers. SGD zigzags in narrow valleys, Momentum builds up velocity, RMSProp scales each dimension's step, Adam combines the two and adds bias correction, and AdamW moves weight decay outside the moment estimates. The slides close with practical advice: Adam(W) is a good default in many cases, and SGD+Momentum can do better but needs more tuning of the learning rate and schedule.
Techniques T16 re-sorts the whole course into three families: how to exploit features (kernels, aggregation, extraction, low-dimensional compression), how to optimize (gradients, equivalent problems, multiple steps), and how to fight overfitting (regularization, validation). It then uses four KDD Cup–winning models to show how the pieces combine in practice. The MOOC was recorded in 2016 and its deep learning stops at pre-training. The Fall 2024 on-campus course filled the gap with 302u (the ReLU family, Xavier/He initialization), 303u (momentum, RMSProp, Adam), a 2020 keynote deck, mlmai.ics, and 1126, an 11-model summary. The Fall 2026 versions of these files are scheduled for week 16 and currently return 404.
Lecture 11 of ML Foundations compares PLA, linear regression, and logistic regression on the same score s = wᵀx. The three differ only in their error functions, and scaled cross-entropy upper-bounds the 0/1 error, so both regressions can do classification. The lecture then turns logistic regression into SGD by computing the gradient on one random example, and builds multiclass classifiers from binary ones with OVA and OVO. Lecture 12 uses a feature transform Φ to turn a circular boundary into a line in Z-space. The price is that computation and d_vc both grow with the dimension, so the advice is: try a linear model first. Practice problems are in Fall 2024 HW4.
Lecture 1 of Machine Learning Techniques turns "which separating line is best?" into an optimization problem. Once you fix the scale so that min yₙ(wᵀxₙ+b) = 1, maximizing the margin is the same as minimizing ½wᵀw, which is a standard QP. Lecture 2 uses Lagrange duality to trade a QP with d̃+1 variables for one with N variables and N+1 constraints, then uses the KKT conditions to recover (b, w) from α. Only the points with αₙ > 0, the support vectors, affect the answer. The dual still contains the inner product zₙᵀzₘ, so the dependence on dimension is not really gone until the kernel lecture.
HW3 has two halves. The four written problems run from Newton's method for logistic regression and a convergence analysis of coordinate descent, through backprop, VJPs, and implicit differentiation, to an information-bottleneck view of what deep networks compress. The notebook has you build a BearTensor computation graph in NumPy, topological-sort backprop, and SGD/Momentum/Adam, then train a red-wine quality regressor with it, plus an optional Muon optimizer. Problems and notebook are public; official solutions and hidden tests are not.
CS189 Spring 2026 covers gradient descent in two lectures. Lec 13 derives the learning-rate limit and the condition number from the Hessian's eigenvalues, then moves through momentum, learning-rate schedules, AdaGrad/RMSProp/Adam, and mini-batch SGD. In Lec 15, Dimakis walks through the same material again, starting from a gradient computed by hand on a small data table. Both slide decks, the recordings, Lec 13's handwritten notes, and Discussions 6 and 7 (with solutions) are all publicly accessible.
HW3 has three parts. The programming part has you build an LP solver by vertex enumeration, stack branch and bound on top of it for integer programs, and formulate three word problems. The written part covers integer programming by hand, the ethics of Amazon's delivery routing, PCA via SVD, and a proof that a Laplace prior equals L1. It is due 10/1, so this guide explains structure and concepts only, with no solutions.
07-380 Lec5 turns the Diet Problem from words into min cᵀx s.t. Ax ⪯ b, then draws it: each constraint is a half-plane, the cost is a direction, and cost contours are perpendicular to c. Push a contour in the −c direction until it last touches the feasible region and you always hit a vertex, so solvers only need the intersections of constraint boundaries. Vertex enumeration checks them all; simplex walks greedily from one vertex to a better neighbor.
07-380 Lec6 adds one constraint to an LP, x ∈ ℤᴺ, and the vertex solution may no longer be an integer. Searching the integer points near the LP solution is not guaranteed to work either. The fix: drop the integer constraint (relaxation) to get an LP lower bound, split on a fractional coordinate into xᵢ ≤ floor and xᵢ ≥ ceil, and keep every subproblem in a priority queue ordered by LP objective. The first all-integer solution popped is optimal.
The course site's schedule splits 07-380's first ten lectures into Reasoning Under Certainty, Optimization and Reasoning Under Uncertainty: prove things with logic and plan with search, then write problems as constrained objectives, and finally let a prior in with MAP and turn to probabilistic models. HW1 tests logic plus search, HW2 planning plus LP graphing, HW3 writing solvers plus PCA and MAP derivations.
Synthesis 1: Tracing how seven weeks form a deliberate knowledge arc from symbolic search to language models, revealing the design philosophy from classical AI to modern ML.
Optimization is not an isolated numerical problem: view SGD spectrally, the magnitude of weight updates determines feature learning; Maximal Update Parameterization transfers LR/init across width, and the critical batch size sets the marginal return of trading compute for convergence.
A model uses loss to know how wrong it is and gradients to know which direction to adjust. Gradient descent repeats three things: compute loss, compute gradients, update parameters. The learning rate controls step size — too large and you overshoot, too small and training takes forever.
Lecture 8 moves from a one-dimensional parabola to vector gradients and compares batch GD, SGD, and mini-batches; the learning rate determines whether updates converge, oscillate, or diverge.
Lecture 11 reads public recipes from MiniCPM, DeepSeek, Qwen, and Llama 3: hold most architectural ratios fixed, sweep learning rate and batch at small scale, then choose model/data allocation with IsoFLOPs. μP helps, but normalization, optimizers, and weight decay can break transfer.
Automatic prompt optimization (APO) has evolved from APE/OPRO to GEPA: replacing sparse rewards with linguistic reflection, winning over GRPO by ~6pp with 4-35x fewer rollouts. Meanwhile, tool descriptions are the overlooked prompt -- small wording changes can shift tool selection rates by 10x, and Anthropic's experiments show Claude self-rewriting tool descriptions outperforms human experts. These two lines are converging: eval-driven automatic optimization is eating hand-tuned prompts.