Table of Contents
🌏 中文版
The official handout is titled Homework 8: Reinforcement Learning and contains written plus programming work. Written sections cover synchronous/asynchronous value iteration, algorithm comparison and selection, a REINFORCE walkthrough, actor-critic/A2C, and empirical questions. Programming implements policy/value networks, n-step returns, policy/value losses, and A2C training for Atari Pong. The ZIP supplies agent.py, environment.py, utils.py, test_runner.py, and requirements.txt; this public bundle contains no reference output, while hidden tests remain on Gradescope.
Separate the MDP components
Write down states, actions, rewards, transitions, discount, and terminal conditions before deriving a Bellman target. Common failures bootstrap through terminals, confuse immediate reward with return, or alter the policy with evaluation data.
First executable action and completion
From the bundle, create the documented environment:
conda create -n HW8 python=3.12
conda activate HW8
pip install -r requirements.txt
Run test_runner.py before Pong training. Completion means passing public checks for network shapes, n-step returns, and losses; starting sustained fixed-seed training; and retaining configurations and reward curves. With no reference output or hidden tests, one high-scoring run is not official validation.
References
Loading...