CMU 07-280 Lecture 13: From Reward Hacking to Auditable AI Scientists
Lecture 13 separates alignment into specification, distribution shift, oversight, and corrigibility, then uses benchmark selection, leakage, and post-hoc selection experiments to show why a final paper cannot audit an autonomous research workflow.