Table of Contents
🌏 中文版
Lectures 5–10 form the algorithmic core. The official agenda covers Policy Gradients, Actor Critic, Value-Based RL, Q-learning in Practice, and two Advanced Policy Gradients lectures. Read them by asking what is estimated, where data comes from, and how bias trades against variance.
Policy-based methods
L5 derives policy gradients from a trajectory objective. Reward-to-go, baselines, and advantages reduce variance without changing the desired objective. L6 introduces actor-critic: a critic supplies the actor's update signal, potentially adding function-approximation bias. Sections 3 and 5 connect and extend these ideas.
HW2 experiments with reward-to-go and neural-network baselines. It is a CPU-first assignment; the homework compute ledger owns the timing and hardware details. For self-study, hold the environment fixed, run three seeds, and retain both individual curves and their mean.
Value-based methods
L7–8 move from Bellman backups to DQN and its stability machinery. L9–10 return to advanced policy-gradient methods. Section 4 places DQN beside SAC; compare their update targets, replay buffers, target networks, and entropy terms.
HW3 implements DQN and SAC. Its starter code spans cheap and expensive environments; see the homework compute ledger for GPU estimates. Validate losses, replay, and evaluation in a small environment first.
Completion check
Explain why policy gradients have high variance, how a critic trades variance for bias, why DQN uses replay and target networks, and what entropy contributes to SAC. If any answer is vague, return to the derivation and smallest experiment before spending more compute.
References
Loading...