Skip to content
所有標籤

#temporal-difference

1 篇文章

CMU 07-280 Lecture 22:不知道 Dynamics 時如何做 Q-learning

第 22 講保留 MDP 骨架,拿掉已知 transition 與 reward 的假設;TD learning 用一步 sample 更新 value,Q-learning 再以 off-policy target 直接學最佳 action values。