Skip to content
系列
29 篇文章

CMU 07-280 完整課程導讀

CMU 07-280 完整課程導讀 系列文章

CMU 07-280 完整課程導讀:搜尋、GPT-2、AlphaZero 為什麼在同一門課?

07-280 是 CMU Spring 2026 首開的 AI+ML 核心:24 講、12 個作業編號,從 heuristic search、CSP 與機器學習一路做到 AlexNet、GPT-2、AlphaZero。教材足以自學,但沒有完整公開錄影、Canvas checkpoint 或 Gradescope 回饋。

CMU 07-280 Lecture 1 導讀:AI、機器學習與表示學習的共同問題

Lecture 1 用 alien autoencoder、AI/ML 範圍與 AI 發展史建立全課座標:智慧系統不是模型清單,而是在不確定下把輸入表示成可計算決策。

CMU 07-280 Lecture 2 導讀:從 UCS、Greedy 到 A* 的 heuristic search

Lecture 2 把搜尋拆成 problem、frontier 與 priority:UCS 看已付成本,Greedy 看估計剩餘成本,A* 用 `f=g+h` 合併兩者;tree 與 graph search 的最優條件並不相同。

CMU 07-280 Lecture 3 導讀:Minimax、Alpha-Beta 與 Expectimax

Lecture 3 把單一路徑改成 contingent plan:minimax 對抗最佳對手,alpha-beta 在不改 root value 下跳過無關分支,expectimax 則用機率取代最壞情況。

CMU 07-280 Lecture 4 導讀:CSP、AC-3 與搜尋順序

Lecture 4 利用 variables、domains、constraints 暴露問題結構,再把 DFS 升級成 backtracking、forward checking、AC-3、MRV 與 LCV;重點是更早證明某些選擇不可能成功。

CMU 07-280 Lecture 5 導讀:用 Loss、Risk 與 ERM 定義機器學習

Lecture 5 把機器學習寫成 `X → Y`、loss、risk 與 empirical risk minimization:訓練集只能提供平均已知損失,真正目標仍是未知分布上的 generalization。

CMU 07-280 Lecture 6 導讀:Decision Trees 如何用 Mutual Information 分裂資料

Lecture 6 從 decision stump 遞迴建樹,用 entropy 衡量 label uncertainty,再以 `I(Y;W)=H(Y)-H(Y|W)`選擇分裂;這是計算可行的 greedy ERM,不是全域最佳樹保證。

CMU 07-280 Lecture 7 導讀:Linear Regression 與 Normal Equation

Lecture 7 把 ERM 套到 linear functions 與 squared loss,從一維 slope 推到矩陣形式 `argmin ||y-Xθ||²`,再在 `XᵀX` 可逆時得到 normal equation。

CMU 07-280 Lecture 8 導讀:Gradient Descent、SGD 與 Learning Rate

Lecture 8 從一維 parabola 推到 vector gradient,再比較 batch GD、SGD 與 mini-batch;learning rate 決定更新是收斂、震盪或發散。

CMU 07-280 Lecture 9:Logistic Regression 如何把分類改寫成機率估計

Lecture 9 不直接預測 0 或 1,而以 sigmoid 建模 P(y=1|x),再用 cross-entropy 與凸最佳化學出參數;多類別版本自然延伸成 softmax regression。

CMU 07-280 Lecture 10:Feature Engineering 與 Regularization 如何交換表達力和穩定性

Lecture 10 先用 φ(x) 讓線性模型表達非線性,再以 train/validation/test 分工、L1/L2 regularization 與 model selection 限制新增自由度造成的過度擬合。

CMU 07-280 Lecture 11:從 Logistic Regression 組出第一個 Neural Network

Lecture 11 把單一 logistic neuron 擴成多層網路:linear layer 產生 z、activation 產生 a,多個 neuron 共同學出 feature transform,再以 loss 和 gradient descent 訓練權重。

CMU 07-280 Lecture 12:Backpropagation 如何重用 Chain Rule

Lecture 12 把神經網路視為 computation graph:forward pass 保存中間量,backward pass 從 loss 開始傳遞 upstream gradient,並用線性層、activation 與 softmax 的局部規則一次算出全部參數梯度。

CMU 07-280 Lecture 13:AI Alignment 從 Reward Hacking 走到可稽核的 AI Scientist

Lecture 13 把 alignment 拆成目標規格、distribution shift、監督與修正能力,並用 autonomous AI scientists 的 benchmark selection、data leakage 與 post-hoc selection 實驗說明:只看最終論文不足以稽核整個研究流程。

CMU 07-280 Lecture 14:CNN 如何把影像的空間結構寫進模型

Lecture 14 以 local connectivity 與 parameter sharing 取代全連接影像模型,從 convolution、stride、padding、pooling 走到 AlexNet、GPU data parallelism、ResNet skip connection 與 BatchNorm。

CMU 07-280 Lecture 15:Pre-training、Transfer Learning 與 Fine-tuning 的分工

Lecture 15 將 pretrained model 拆成 representation g 與 task head h:可以凍結 g 只訓練 head,也能用較小 learning rate fine-tune 部分或全部參數;選擇取決於資料量與 source-target 差距。

CMU 07-280 Lecture 16:Maximum Likelihood 如何統一 Logistic 與 Linear Regression

Lecture 16 從 likelihood p(D|θ) 出發,以 i.i.d. 將聯合機率寫成乘積,再用 log 變成和;Bernoulli MLE 得到樣本比例,conditional Bernoulli 得到 logistic cross-entropy,Gaussian noise 則得到 squared error。

CMU 07-280 Lecture 17:從 Tokenization 到 N-gram Language Model

第 17 講先決定文字如何切成 token,再用 N-gram 把序列機率改寫成可從 corpus 計數的條件機率;tokenization 不是前處理小事,而是模型能看見什麼的第一個設計決定。

CMU 07-280 Lecture 18:N-gram 如何訓練、取樣與失敗

第 18 講把 chain rule 截成 N-gram Markov assumption,以 corpus counts 做 MLE,再比較 greedy、categorical sampling 與 temperature;真正的瓶頸是未見 context 的零機率與固定視窗。

CMU 07-280 Lecture 19:Word Embedding 如何把下一詞預測變成幾何

第 19 講以兩個 embedding matrices、dot-product similarity、softmax 與 cross-entropy 建立最小 next-token model,讓相似 context 透過共享向量參數取代 N-gram 的獨立計數格。

CMU 07-280 Lecture 20:從 Position Encoding 推到 Causal Self-Attention

第 20 講先把單 token embedding 擴成 sequence,以 positional encoding 補順序,再推導 Q/K/V scaled dot-product attention、causal mask 與 multi-head blocks,最後接到 GPT-2。

CMU 07-280 Lecture 21:Bellman Equation 如何解 Markov Decision Process

第 21 講把隨機序列決策寫成已知 dynamics 的 MDP,以 Bellman backup 定義 value 與 Q-value,再用 value iteration 或 policy iteration 求最佳 policy。

CMU 07-280 Lecture 22:不知道 Dynamics 時如何做 Q-learning

第 22 講保留 MDP 骨架,拿掉已知 transition 與 reward 的假設;TD learning 用一步 sample 更新 value,Q-learning 再以 off-policy target 直接學最佳 action values。

CMU 07-280 Lecture 23:從 Approximate Q-learning 到 DQN

第 23 講以 Qθ(s,a) 取代巨大 Q-table:先用 features 線性近似並由 squared TD error 推導 gradient update,再以 replay data 與固定 target network 形成 DQN。

CMU 07-280 Lecture 24:Monte Carlo Tree Search 如何接上 AlphaZero

Spring 2026 Lecture 24 是 MCTS,不是 Fall 2026 的 LLM post-training;本講以 selection、expansion、rollout、backup 與 UCB 分配模擬,再由 policy/value heads 與 self-play 接到 AlphaZero。

CMU 07-280 階段複習一:從搜尋問題走到監督式學習

第一階段把 Lectures 1–12 串成同一條決策鏈:先定義狀態、動作與目標,再用 heuristic、loss、regularization 與 backpropagation 控制龐大搜尋空間。

CMU 07-280 階段複習二:用 AlexNet 與 GPT-2 把模型真正組起來

第二階段不把 CNN 與 Transformer 當兩份架構圖背誦,而是透過 HW8 與 HW11 檢查表示、計算圖、訓練、遷移與生成是否真的接得起來。

CMU 07-280 階段複習三:從 MDP、Q-learning 到 AlphaZero

第三階段把 value、policy、bootstrapping、function approximation 與 MCTS 接成 AlphaZero:network 提供先驗與估值,search 改善決策,self-play 再產生下一輪資料。

CMU 07-280 全課總結:學會什麼、缺什麼,以及下一門怎麼選

完成 07-280 不等於看完 24 篇導讀;至少要留下搜尋器、監督式模型、CNN/GPT-2 實驗與一個 RL+MCTS 小系統,再依缺口選 07-380、10-301 或專題課。