🌏 中文版
Lecture 8 of CMU 07-380 AI & ML II, MAP, met on Monday, September 21, 2026. It opens the "Reasoning Under Uncertainty" block of the schedule. The previous three lectures solved deterministic optimization problems (LP, ILP, PCA). This one brings probability back in with a practical question: when data is scarce, what else do we know? The answer is a prior, and the method that folds a prior into estimation is maximum a posteriori (MAP) estimation.
This guide reflects the course site as of 2026-09-29. The schedule is marked subject to change.
Official materials and what I read
Materials I opened and read:
- Lec8 slides (pdf), plus pptx and inked versions. Instructors: Pat Virtue and Mohammad Salameh.
- Pre-reading PR5 MAP.pdf, with two worked examples: a trick coin and course hours.
- Desmos: Bernoulli Likelihood and Posterior, where you can adjust the prior and the flips yourself.
- Recitation 5 and its solutions (Friday 9/25, the same session as Quiz 2).
- The 07-280 background notes the site assigns: Notation Guide, Probability Background, and MLE Pre-reading. Further reading: MML 9.2.3–9.2.4 and Mitchell's Estimating Probabilities.
A note on limits. The slides' text layer holds only titles, equation skeletons, and poll questions. Most of the derivation on the "Regularization and MAP" slides was handwritten in class. The ink in the inked version is image data, and I did not transcribe it. So when I say "the slides" below, I mean only the printed text. The pre-reading checkpoint lives on Canvas and is CMU-only, and the site links no recordings. On this site's A0–A3 scale, this lecture's materials (slides, notes, recitation with solutions) reach A3. The course as a whole is still A2 while it runs.
The question carried forward: MLE struggles with little data
I won't re-derive MLE here. See the 07-280 Lecture 16 MLE guide. The one thing to remember: MLE picks the θ that maximizes p(D|θ). It looks only at the data.
The MLE/MAP comparison table at the start of PR5 states the problem directly. Without prior assumptions, we know nothing about the parameters before collecting data, and with little data we risk overfitting. MAP takes the opposite stance. Even with zero data you have an initial estimate, at the price of one more modeling assumption (the prior). If that prior is reasonably accurate, a small amount of data already gives a good estimate.
Bayes rule has two uses; this lecture needs the first
The slides put two versions of Bayes rule side by side:
p(θ|D) = p(D|θ) p(θ) / p(D) ← data and parameters: this lecture
p(y|x) = p(x|y) p(y) / p(x) ← inputs and outputs: next lecture's generative models
In the first, p(θ|D) is the posterior, p(D|θ) the likelihood, and p(θ) the prior. When we search for the best parameters, p(D) does not depend on θ, so we drop it:
p(θ|D) ∝ p(D|θ) p(θ)
A Recitation 5 concept question asks exactly this: why can MAP ignore p(D)? Because we take the argmax over θ, and p(D) is a constant.
The second form is not needed until next time. Showing it now makes the point that one formula plays two roles. The Lecture 9 guide on generative models picks up from there.
Worked example 1: the trick-coin posterior tables
PR5's story: you buy a "randomly weighted trick coin" at a joke shop. Before you flip it, you find the shop's invoice in a trash bin. It lists five coin types and their quantities. As a prior, the heads probability ϕ can take only five values:
| ϕ | 0.0 | 0.2 | 0.5 | 0.8 | 1.0 |
|---|---|---|---|---|---|
| p(ϕ) | 0.20 | 0.25 | 0.40 | 0.05 | 0.10 |
Before any flip, the MAP estimate is ϕ = 0.5, the largest prior (the notes give 80/200 = 0.4). Each flip multiplies the likelihood by one more factor: ϕ for heads, (1−ϕ) for tails. Finally, divide by the sum Z over all ϕ values to normalize.
This snippet reproduces the notes' tables for N = 0, 1, and 5:
prior = {0.0: 0.20, 0.2: 0.25, 0.5: 0.40, 0.8: 0.05, 1.0: 0.10}
def posterior(flips):
unnorm = {}
for phi, p in prior.items():
like = 1.0
for f in flips:
like *= phi if f == "H" else (1 - phi)
unnorm[phi] = p * like
Z = sum(unnorm.values())
return {phi: v / Z for phi, v in unnorm.items()}, Z
for flips in ["", "H", "HTTTT"]:
post, Z = posterior(flips)
print(flips or "{}", round(Z, 6), {k: round(v, 6) for k, v in post.items()})
For HTTTT you get Z = 0.033044 and a posterior peak at ϕ = 0.2 (0.619780), matching the notes. Three things to notice:
- The first heads zeroes out ϕ = 0.0 immediately, because the likelihood picks up a factor of 0.
- After a single heads, MAP is still 0.5 (0.512821). One data point cannot outweigh the prior.
- Four more tails move MAP to 0.2. More data means more likelihood factors and less influence from the prior.
Poll 4 on the slides asks about that third point: as data grows, what happens to MAP, and which distribution does the posterior approach? The Recitation 5 solutions reach the same conclusion. The prior's relative influence shrinks, and MLE and MAP converge to the same value.
The estimation recipe: MAP adds one term
The slides and PR5 both write the two methods as four steps. Only step 1 differs:
| Step | MLE | MAP |
|---|---|---|
| 1 | Write the likelihood `p(D | θ)` |
| 2 | `J(θ) = −log p(D | θ)` |
| 3 | Compute ∂J/∂θ | Same |
| 4 | Set the derivative to zero and solve, or use (stochastic) gradient descent | Same |
Expanding step 2 shows the key idea:
J(θ) = −log p(D|θ) − log p(θ)
The first term is the original loss. The second is the extra term contributed by the prior. Everything in the regularization section below starts from this line.
With only five candidate values, the trick coin can be solved by computing every posterior and picking the largest. With continuous parameters you cannot enumerate, which is why steps 3 and 4 need derivatives or gradient descent.
Worked example 2: a Beta prior and an ad click rate
Next the slides introduce the Beta distribution:
p(ϕ; α, β) = ϕ^(α−1) (1−ϕ)^(β−1) / B(α, β)
A Bernoulli likelihood times a Beta prior gives a Beta posterior: Beta(α + N_{y=1}, β + N_{y=0}). When prior and posterior share a family, the prior is called conjugate. The slides list two more pairs: Categorical with Dirichlet, and Gaussian with Gaussian. Their mnemonic: think of Beta as having already seen α−1 heads and β−1 tails. They also note the special case: with a uniform prior, MLE and MAP are identical.
Recitation 5, problem 2, turns this into a calculation. An ad is shown to N people and N₁ click. MLE gives N₁/N. With a Beta(α, β) prior, the negative log posterior simplifies to:
argmin −(N₁ + α − 1) log ϕ − (N − N₁ + β − 1) log(1 − ϕ)
Setting the derivative to zero:
ϕ̂_MAP = (N₁ + α − 1) / (N + α + β − 2)
The problem uses N = 100 and N₁ = 10, and sets α = 7, β = 95 based on "ads usually get about 6% clicks." So MAP = 16/200 = 0.08, while MLE is 0.10. The solutions do not declare a winner. If this ad resembles the past ads that averaged 6%, MAP is more credible. If it ran under different conditions, MLE may be the better choice. A prior is an assumption you have to defend.
Open the Desmos Bernoulli posterior and move α, β, and the flip counts. It builds intuition faster than the formulas do.
Worked example 3: a Gaussian prior and course hours
PR5 section 5 is the continuous version. You ask four classmates how many hours per week a systems course takes and get D = {18, 20, 14, 10}. Past course evaluations report a mean of ν = 23.9 with standard deviation τ = 1.56, which becomes a Gaussian prior on μ. Following the recipe gives:
μ̂_MAP = (σ² ν + τ² Σx⁽ⁱ⁾) / (σ² + N τ²)
The notes then make a "rather naive" assumption, σ = τ, and the formula collapses to treating the prior mean as one extra data point: (ν + Σx⁽ⁱ⁾) / (1 + N). This is the same intuition as Beta's imaginary flips.
If you check the arithmetic, you'll find a small typo. The notes write (23.9 + 18 + 20 + 14 + 10) / 5 = 17.8, but 85.9 / 5 is 17.18. The takeaway holds: MLE is 15.5, and the prior pulls the estimate up toward 23.9.
Linear regression: MAP is regularization
The second half of the slides returns to linear regression. First the probabilistic reading: assume y ~ N(wᵀx + b, σ²), and conditional MLE becomes minimizing squared error (derived in 07-280; see Lecture 7 on linear regression and Lecture 16). Then the slides ask: for regression with polynomial features, what do we want to assume about the parameters? They answer by placing a Gaussian prior on the weights.
The small "Regularization and MAP" table in Recitation 5 states the mapping most clearly:
| Regularization | Penalty | Equivalent prior |
|---|---|---|
| Ridge regression | ‖w‖₂² | w_j ~ N(0, τ²) |
| Lasso | ‖w‖₁ | w_j ~ Laplace(0, b) |
Why does this hold? Go back to step 2: J = −log p(D|w) − log p(w). Below is my own derivation of the Gaussian-prior case following the recipe (these slides were handwritten, so I did not copy the in-class derivation):
−log p(D|w) = (1/2σ²) Σ (y⁽ⁱ⁾ − wᵀx⁽ⁱ⁾)² + const
−log p(w) = (1/2τ²) ‖w‖₂² + const
Multiply by 2σ²: Σ (y⁽ⁱ⁾ − wᵀx⁽ⁱ⁾)² + (σ²/τ²) ‖w‖₂²
So ridge's λ is σ²/τ². A narrower prior (smaller τ) means a larger λ and weights pulled harder toward zero. This is the probabilistic origin of 07-280 Lecture 10 on regularization. You think you are tuning a penalty coefficient, but you are really choosing a prior.
The Laplace-prior-to-L1 derivation is yours to do in HW3 written problem 4, so I leave it out.
What else Recitation 5 covers
The first three sections of Recitation 5 are about MAP: definitions, the click-rate problem, and three concept questions (why ignore p(D), how the two estimates change with more data, and when they coincide). The answer to the third is worth remembering. With a uniform prior the two are equal, so MLE can be viewed as a special case of MAP.
Section 5 has a 2×2 table that sorts the models seen so far by discriminative/generative and MLE/MAP:
| MLE | MAP | |
|---|---|---|
| Discriminative | Linear regression, logistic regression, logistic regression with polynomial features | Linear regression with L2, logistic regression with a Laplace prior |
| Generative | Naive Bayes | Naive Bayes with Laplace smoothing |
The second half, Gaussian Discriminant Analysis and Naive Bayes, belongs to the next lecture, and I cover it in the Lecture 9 guide. One easy confusion to flag now: the solutions note that the "Laplace" in Laplace smoothing has nothing to do with the Laplace distribution. For a Bernoulli, Laplace smoothing is equivalent to a Beta prior.
Related reading
- The full MLE derivation: 07-280 Lecture 16.
- Regularization from the loss-function side: 07-280 Lecture 10.
- The previous post is Lecture 7 on PCA and low-rank optimization. The next is the HW3 guide, which puts LP, ILP, PCA, and this lecture's MAP into one assignment.
Things to do tonight
- Run the trick-coin snippet above with the sequence
HHHHHand watch when MAP jumps from 0.5 to 1.0. - Work the Recitation 5 click-rate problem by hand, then multiply both α and β by 10 (a more confident prior) and see how far MAP moves toward 0.06.
- Derive Gaussian prior = L2 yourself and confirm λ = σ²/τ². Then do the Laplace version in HW3.
References
- CMU 07-380 Fall 2026 course site (Schedule, Recitation, Policies)
- 07-380 Lec8 MAP slides (pdf)
- 07-380 Pre-reading: Maximum a posteriori Estimation (PR5)
- 07-380 Recitation 5 and Recitation 5 Solutions
- Desmos: Bernoulli Likelihood and Posterior
- 07-280 Notation Guide, 07-280 Probability Background, 07-280 MLE Pre-reading
- Tom Mitchell, Estimating Probabilities: MLE and MAP
- Deisenroth, Faisal, Ong, Mathematics for Machine Learning
- CMU 07-380 Fall 2026 Overview
Loading...