CMU 07-280 Lecture 13:AI Alignment 從 Reward Hacking 走到可稽核的 AI Scientist
Lecture 13 把 alignment 拆成目標規格、distribution shift、監督與修正能力,並用 autonomous AI scientists 的 benchmark selection、data leakage 與 post-hoc selection 實驗說明:只看最終論文不足以稽核整個研究流程。
Lecture 13 把 alignment 拆成目標規格、distribution shift、監督與修正能力,並用 autonomous AI scientists 的 benchmark selection、data leakage 與 post-hoc selection 實驗說明:只看最終論文不足以稽核整個研究流程。
第 16 講把 NLP 的社會影響拆成四題:模型為何 hallucinate、AI 輔助創作的同質化悖論、工作如何重組,以及價值對齊為何不能化約成單一 reward。
Patronus AI 把 evaluator 當成可重用評分單元,再接到離線 experiment 與線上 trace;它適合要現成幻覺、安全與多模態評估模型的團隊,但 judge 分數不能取代人工標註、確定性測試或真正的資安驗證。