ai deep-dive Stanford CS229 導讀 2026年8月22日 強化學習:MDP、價值迭代與連續狀態 第 19 章用 Bellman 方程把長期決策拆成一步更新,並從已知 MDP 的 value iteration 走到模型學習與連續狀態近似。 #cs229#reinforcement-learning#mdp#value-iteration#fitted-value-iteration