🌏 中文版
This guide follows the public materials of CS189 Spring 2026 (Jennifer Listgarten / Alex Dimakis). The series starts at the Berkeley CS189 overview.
The previous post was about how to minimize a loss. This one asks where those losses come from. Two things you have already seen come back from a new angle:
- Ridge regression: introduced as least squares plus
λ‖w‖². Lec 14 shows it is exactly the MAP estimate under a Gaussian prior on the weights. - The logistic regression loss: introduced as maximum likelihood. Lec 16 recasts it as cross-entropy, a measure in bits of how far apart two distributions are.
The midterm (3/17) comes the week after these lectures, so the end of this post lays out a self-check using the official exam.
Where the materials are
| Item | Official title / content | Materials | Assigned Bishop reading |
|---|---|---|---|
| Lec 14 (3/5) | MLE, MAP and Bias-Variance Trade-off | Notes folder: lec14.pdf (58 pages); recording | 2.6.1–2.6.2, 3.1.1, 4.1.2, 4.1.6, 4.3 (bias-variance), 5.4.3 |
| Lec 16 (3/12) | Entropy, Information and Logistic Regression | Notes folder: lec16.pdf (66 pages); recording | 2.5.1 (entropy), 2.5.5 (KL), 5.3.1, 5.4.3–5.4.4 |
| Midterm (3/17, 7–9pm) | 6 problems + Honor Code, 56 points, 110 minutes | Exam sp26-midterm.pdf, solutions sp26-midterm-sol.pdf, walkthrough playlist (6 videos, Problems 1–6) | — |
Both midterm PDFs sit in the past-exams folder linked from the Resources page (under midterm/exams and midterm/solutions). The same folder also has the Fall 2025 midterm and practice midterm. I opened all of these links without logging in on 2026-09-29. On this site's A0–A3 scale, this stretch is A3, with the bonus of a real exam that comes with solutions. What you can't get is the in-class Slido interaction and the midterm grade distribution.
Lec 14: MLE, MAP, and ridge
The roadmap in lec14.pdf has five parts: least squares as maximum likelihood (a recap), choosing different noise models, prior beliefs, MLE vs MAP, and bias-variance.
It starts with a coin. The probability Y of heads is unknown, and you observe two outcomes x1 and x2. MLE looks only at the data. MAP first puts a prior on Y (the slides use a discrete prior and mention that a Beta prior also works), then multiplies by the likelihood using Bayes' rule. The difference: with little data, the prior pulls the estimate toward what you already believed.
Then it moves to regression weights. The same idea applies to w: put a Gaussian prior centered at 0 on it, with variance σ_w² controlling how strongly you expect small weights. The slides read this prior as a preference for simpler models. Taking the negative log of the posterior turns maximization into minimization, and after rearranging you get exactly the ridge objective. The slides conclude:
- For linear regression with Gaussian noise, least squares equals MLE.
- Add a Gaussian prior on w, and ridge equals MAP.
Where the regularization strength λ comes from
After taking the negative log, the data term is weighted by 1/σ² (the noise variance) and the prior term by 1/σ_w² (the prior variance on the weights). The slides multiply the whole expression by σ², so λ is the ratio of the two variances. More noise, or a stronger belief that weights should be small, means a larger λ.
Lec 14: bias-variance
The slides first list what a model is supposed to do (fit the data, explain what we observe, generalize, predict the future, and so on), then use "Is this cat grumpy, or are we overfitting to human faces?" to introduce overfitting. Next they split the expected test error into three terms:
| Term | Definition in the slides | Associated with |
|---|---|---|
| Bias | The expected deviation between the predicted value and the true value; depends on the choice of function class | Underfitting |
| Noise | Randomness in the data-generating process itself: measurement variability, stochasticity, missing information | Beyond your control |
| Model variance | How much the prediction changes across different training datasets | Overfitting |
The key trick in the derivation is adding and subtracting a term: insert h(x) into the error expression, then use the independence of the noise ε and w to make the cross terms vanish. The slides include two quick quizzes, such as "what is E[t] equal to?" and "what is E[ε] equal to?", that you can use to check you're following.
Apply this to ridge. As λ grows, the weights shrink and the model becomes less flexible, so bias goes up, but the model is less sensitive to noise, so variance goes down. As λ approaches 0, the opposite happens. The slides sum it up in one line: regularization is a mechanism that trades variance for bias. The experiment follows Bishop's setup: N = 25 noisy observations, fit with M = 24 Gaussian basis functions plus a bias, repeated over many generated datasets. Training error rises monotonically with λ, test error is U-shaped, and λ is chosen by validation.
End of Lec 14: Chatbot Arena and how to read a paper
The last part of the deck introduces Chatbot Arena, a public platform where a user enters a prompt, two anonymous models answer side by side, the user votes for the better one, and many such battles are aggregated into a leaderboard. The slides tie it back to this lecture: leaderboard scores come from a Bradley–Terry model, which is essentially logistic regression (Y = which model won), and a model's Arena Score is its coefficient β.
Then comes a checklist of five questions for reading a paper:
- What problem is this paper tackling?
- What do prior works do, and where do they fall short?
- What is the key insight? What does it do that prior works don't?
- What are the inputs and outputs of the method?
- What are its limitations?
The slides note that the first four can usually be found in the introduction. This section maps directly onto the HW2 paper questions, so read these pages before starting the homework.
Lec 16: entropy through compression
lec16.pdf doesn't open with a formula. It opens with a compression problem. Suppose it rains in Seattle each day with probability 80%, independently, and you record 100 days. How much information does that sequence contain, and how far can you compress it?
- Entropy: Shannon's compression theorem says entropy is the ultimate limit of compression. In this example, 100 bits compress to about 73 bits. If it always rains, there is no information and you can compress to 0 bits. A fair coin has entropy of 1 bit per flip and can't be compressed at all.
- Intuition: pointing out one item among L takes log L bits. Random sequences almost always land in a set of "typical sequences", so you only need log(number of typical sequences) bits. The slides say this is essentially the law of large numbers (the Asymptotic Equipartition Property) and recommend Cover & Thomas, Elements of Information Theory, for details.
- Exercise: a horse race with 8 horses. A naive code needs 3 bits. The slides ask you to compute the entropy and check whether a given code reaches the Shannon limit.
Lec 16: KL, cross-entropy, and back to logistic regression
The slides then reinterpret two quantities in terms of bits:
- Cross-entropy: the total number of bits needed to describe data whose true distribution is p when you use a compression scheme designed for q.
- KL divergence: the extra bits wasted compared with the optimal scheme.
So cross-entropy = entropy + KL. The slides put it in one line: you can't compress p using a scheme designed for q. They also preview that cross-entropy will be the loss used for both logistic regression and deep learning.
Back to logistic regression, the slides first stress that "logistic regression is not a regression; it's binary classification", and contrast it with generative models: assume x is Gaussian in each class with a shared covariance and you get LDA; with different covariances, QDA. A worked example then predicts whether a customer clicks a home insurance ad from two features, with parameters w = [0.1, 1, 10]. It computes the score, passes it through a sigmoid to get a click probability, writes the likelihood of the whole dataset, and shows that the log-likelihood is the negative cross-entropy.
For multiple classes, the sigmoid becomes a softmax, the weights become a matrix, probabilities are softmax(Wx), and the loss is −Σ yᵢ log pᵢ. The slides use MNIST (60,000 handwritten digits) and CIFAR-10 (6,000 images per class, 50,000 for training and 10,000 for testing) as example datasets.
Midterm: self-assessing with the official exam
The exam is 18 pages with 6 problems, and the 1-point Honor Code brings the total to 56. After reading the problems, here is how each maps back to the lectures:
| Problem (original title) | Points | What it tests | Review |
|---|---|---|---|
| A Linear Affair | 5 | Multiple-choice concepts on linear regression: linear in what, and what happens when n ≫ d or n ≪ d | Lec 7–10 |
| Smooth Operator | 7 | A smoothness regularizer penalizing differences between adjacent weights: write it as ‖Dw‖², find the closed form, examine the λ → ∞ limit | Lec 9–10 |
| Chill Gaussian Question | 11 | MLE and MAP (Gaussian prior) for the mean of a multivariate Gaussian, and how MAP behaves as n → ∞ | Lec 5–6, 14 |
| Ozan Risks It All for MoG | 11 | GMM log-likelihood, responsibilities rᵢₖ, and the reduction to k-means as σ² → 0 | Lec 4–7 |
| We Adopt a Sigma Grindset | 10 | Which of LDA / QDA / logistic regression is generative vs discriminative; with a shared covariance the log-odds reduce to a linear form, and how LDA relates to logistic regression; bias and variance of LDA / QDA when covariances differ | Lec 11–12, 14, 16 |
| On a Downward Spiral | 11 | A gradient and one update step computed by hand on three data points, properties of SGD gradients, what happens to the gradient when one feature is scaled by 100, and which model has no closed form | Lec 13, 15 |
I didn't find entropy, KL, momentum, or Adam in the problems; Lec 16 shows up mainly through the logistic regression problem. That describes this one exam, though, not the official exam scope.
A suggested self-check:
- Print
sp26-midterm.pdf, set a 110-minute timer, and finish it without notes. - Grade yourself against
sp26-midterm-sol.pdf(25 pages), and use the table above to tag each lost point with a lecture. - Watch only the walkthrough videos for the problems you missed. Note the playlist order is Problem 1, 2, 3, 4, 6, 5, so Problem 5 is last.
- For another round, the same folder has the Fall 2025 midterm and practice midterm, with solutions. A different set of instructors wrote those, which helps you check whether you've only learned one exam style.
Matching Fall 2026 lectures
Fall 2026 moves bias-variance earlier, to Lecture 7 "Bias-Variance Trade-off + Regularization", with Logistic Regression in Lectures 8–9. The midterm is on 10/20 (Week 9), preceded by a Midterm Review on 10/16. The Fall 2026 schedule has no lecture titled for entropy or KL.
Further reading
- Probability background: Stanford CS109 on the Beta distribution, information theory, maximum likelihood estimation, logistic regression
- Classic ML for comparison: Stanford CS229 guide
Previous: Lec 13 & 15: convergence, momentum, Adam, SGD. Next: HW2 guide.
Things you can do tonight
- Derive the ridge objective yourself from the negative log posterior, and write down which two variances λ is the ratio of.
- Compute the entropy H(0.8) for the Seattle example and confirm that 100 days come to roughly 73 bits.
- Pick an evening and sit the sp26 midterm under time, following the steps above.
References
- CS189 Spring 2026 homepage and schedule
- CS189 Spring 2026 Resources (entry point to past exams)
- Lecture 14 Notes folder (lec14.pdf)
- Lecture 14 recording: MLE, MAP and Bias-Variance Trade-off
- Lecture 16 Notes folder (lec16.pdf)
- Lecture 16 recording: Entropy, Information, and Logistic Regression
- CS189 Spring 2026 midterm (sp26-midterm.pdf)
- CS189 Spring 2026 midterm solutions (sp26-midterm-sol.pdf)
- Midterm walkthrough playlist
- Chiang et al., Chatbot Arena (arXiv:2403.04132)
- CS189 Fall 2026 schedule
Loading...