Skip to content
All tags

#emergent-misalignment

2 posts

CS2881R HW0: Reproducing Emergent Misalignment with a 1B Model

CS 2881R's HW0 was the admission filter: LoRA-fine-tune Llama-3.2-1B-Instruct on bad medical, financial, or extreme-sports advice, then check whether it turns harmful on unrelated questions too. The repo ships encrypted training data, generate.py, and a judge.py that uses gpt-4o-mini as grader; the README targets alignment below 75 and coherence above 50. train.py is empty and yours to write. For self-study, know three things: the grading script only prints averages and never decides pass/fail, refusals drop out of the average, and the base-model baseline is 20 medical questions while your CSV is 10 medical plus 10 non-medical.

CS2881R L1: Why AI Safety Deserves a Graduate Course

CS 2881R's first lecture (2025-09-04) opens with three pre-readings. AI 2027 sketches recursive self-improvement reaching superhuman AI within five years. AI as Normal Technology argues AI will diffuse slowly, like electricity. METR measures the length of human tasks an AI can finish half the time and finds it doubling about every 7 months. Boaz's lecture splits AGI definitions into capability-based and impact-based, and alignment approaches into principles, character training, and model specs. The student experiment runs HW0 in reverse: fine-tuning on aligned bioethics answers also raised alignment scores on environmental-policy questions.