CMU 07-280 Stage Review III: From MDPs and Q-learning to AlphaZero
Stage III connects value, policy, bootstrapping, function approximation, and MCTS into AlphaZero: a network supplies priors and estimates, search improves decisions, and self-play creates the next training set.