Skip to content
所有標籤

#fitted-value-iteration

1 篇文章

強化學習:MDP、價值迭代與連續狀態

第 19 章用 Bellman 方程把長期決策拆成一步更新,並從已知 MDP 的 value iteration 走到模型學習與連續狀態近似。