Skip to content
系列
23 篇文章

Stanford CS109 導讀

逐講讀 Stanford CS109:機率、隨機變數、推論與模擬,補齊機器學習與資料科學真正會用到的機率底座。

Stanford CS109 導讀:一門機率課把「怎麼用語言模型讀這一講」寫成了官方教材

CS109 在 2026 年夏季的每一講旁邊,掛了一份官方寫的 LLM Learning Guide——六個概念、每個概念一組 Learn 與 Test me 提示詞,逐週產出共 23 份 PDF。同一門課的榮譽守則第 4 條卻明文禁止拿 LLM 解作業,而成績有 65% 壓在現場考試。這兩件事是同一套設計的兩半。

Stanford CS109 Lecture 1|What is Probability?:先把隨機問題列成結果集合,再談事件的機率。

先把隨機問題列成結果集合,再談事件的機率。

Stanford CS109 Lecture 2|Conditional Probability:條件不是裝飾,而是把樣本空間縮到已知資訊仍允許的部分。

條件不是裝飾,而是把樣本空間縮到已知資訊仍允許的部分。

Stanford CS109 Lecture 3|Bayes Theorem:Bayes 定理把容易建模的生成方向,翻成真正想問的推論方向。

Bayes 定理把容易建模的生成方向,翻成真正想問的推論方向。

Stanford CS109 Lecture 4|Counting and Combinatorics:先判斷順序是否重要、元素能否重複,公式才不會套錯。

先判斷順序是否重要、元素能否重複,公式才不會套錯。

Stanford CS109 Lecture 5|Random Variables and Expectation:隨機變數是把結果映成數字;期望值是加權平均,不保證會真的出現。

隨機變數是把結果映成數字;期望值是加權平均,不保證會真的出現。

Stanford CS109 Lecture 6|Moments:期望值、LOTUS 與線性性

期望值把分布壓成加權平均;LOTUS 處理變換後的值,linearity 則讓隨機變數的和即使不獨立也能直接計算。

Stanford CS109 Lecture 7|Variance 與 Poisson:從分散程度到稀有事件計數

先用 variance 描述隨機變數的分散程度,再用 Poisson 處理固定區間內的事件數,以及大 n、小 p 的二項近似。

Stanford CS109 Lecture 8|Continuous Random Variables:PDF、CDF、Uniform 與 Exponential

連續變數的單點機率為零,區間機率是 PDF 面積;CDF、Uniform 與 Exponential 則把面積、等待時間與 memorylessness 串起來。

Stanford CS109 Lecture 9|Normal Distribution:標準化、Φ 與 continuity correction

標準化把不同尺度的 Normal 變數轉成 Z;Φ、線性轉換與 continuity correction 再把區間與大型 binomial 變成可計算的機率。

Stanford CS109 Lecture 10|Probabilistic Models:joint、marginal、independence 與 Bayes

Joint distribution 保存多個變數的完整關係;marginal、conditional、independence 與 Bayes 都是從這份關係表取出不同問題的答案。

Stanford CS109 Lecture 11|Inference:prior × likelihood → normalize

Inference 把 hidden variable 的 prior 逐項乘上 observation likelihood,再正規化成 posterior;同一迴圈可處理多次觀察與離散化的連續 belief。

Stanford CS109 Lecture 12|General Inference:Bayesian networks、sampling 與 rare evidence

Bayesian network 用 conditional independence 分解巨大 joint;ancestral sampling 生成 joint samples,rejection sampling 再以 evidence 篩出 conditional。

Stanford CS109 Lecture 13|Multinomial:多類別計數、bag of words 與 log probability

Multinomial 把 binomial 的兩類計數推廣到多類;同一 PMF 也能把文件視為 word counts,配 Bayes 與 log-score 做 authorship inference。

Stanford CS109 Lecture 14|Beta:把未知 probability 變成可更新的 random variable

Beta distribution 表示對未知成功率的完整 belief;success/failure data 只需更新兩個參數,便能取得 posterior、平滑估計與 Thompson-sampling decision。

Stanford CS109 Lecture 15|Adding Random Variables 與 Central Limit Theorem

少數分布的 independent sums 有 closed form;一般 IID sums 則由 CLT 在大樣本下近似 Normal,離散 sums 還需 continuity correction。

Stanford CS109 Lecture 16|Bootstrapping:sampling statistics、error bars 與 p-values

Bootstrap 把 sample histogram 當作 population proxy,以 replacement 重抽並重算 statistic,近似 estimator 的 sampling distribution、error bar 與 null p-value。

Stanford CS109 Lecture 17|Algorithmic Analysis:conditional expectation、indicators 與 recursion

隨機程式的 expected cost 可依第一個 random choice 分情境;計數問題則拆成 indicators,兩者都靠 linearity,而不必硬求完整 distribution。

Stanford CS109 Lecture 18|Information Theory:surprise、entropy、information gain 與 KL

Surprise 把低機率事件轉成 bits;entropy 是 expected surprise,information gain 選擇最能降低 uncertainty 的問題,KL 則量化錯用 model distribution 的額外代價。

Stanford CS109 Lecture 19|Maximum Likelihood Estimation:固定資料,找最能解釋資料的參數

MLE 固定觀察資料、最佳化參數;log-likelihood 讓乘積變加總,但最大值也可能落在邊界。

Stanford CS109 Lecture 20|Logistic Regression:從 Bernoulli likelihood 推出 gradient

Logistic regression 用 sigmoid 把線性分數變成 Bernoulli 機率,而 gradient xⱼ(y-ŷ) 直接來自 log-likelihood 的 chain rule。

Stanford CS109 Lecture 21|Comparing Classifiers:accuracy 之外還要問 calibration、錯誤成本與 fairness

比較 classifier 不能只看 accuracy;還要用 held-out data、baseline、calibration、precision/recall 與明確的 fairness criterion。

Stanford CS109 Lecture 22|Deep Learning:用 chain rule 推出 backpropagation

Neural network 是堆疊的 logistic units;forward pass 算機率,backpropagation 重用 output error 來計算所有 gradients。