Skip to content
Series
16 posts

Reading Harvard CS2881R

A lecture-by-lecture reading of Harvard CS 2881R AI Safety (Boaz Barak, Fall 2025). It starts from the emergent-misalignment HW0, then covers safety training, jailbreaks and prompt injection, model specs and content policies, scheming and interpretability, recursive self-improvement, capability measurement, and the economic and mental-health impacts, ending with the students' reproduction and final research projects. It is based on the public recordings, reading lists, slides and assignment specs.

Reading Harvard CS2881R: What Outsiders Can Get from the First Graduate AI Safety Course

Harvard CS 2881R is the graduate AI safety seminar Boaz Barak first taught in Fall 2025. That term is finished: all 12 reading lists are public, the YouTube playlist has lecture recordings for 11 of the 12 sessions, HW0 is a GitHub repo you can run yourself, and the midterm and final specs and rubrics are out. This series rates it A3 by seminar standards. The gaps are just as clear: no traditional problem sets, slides for only about half the sessions, no lecture recording for L5, and only the opening remarks for L8. Fall 2026 is in progress and is treated only as a preview.

CS2881R L1: Why AI Safety Deserves a Graduate Course

CS 2881R's first lecture (2025-09-04) opens with three pre-readings. AI 2027 sketches recursive self-improvement reaching superhuman AI within five years. AI as Normal Technology argues AI will diffuse slowly, like electricity. METR measures the length of human tasks an AI can finish half the time and finds it doubling about every 7 months. Boaz's lecture splits AGI definitions into capability-based and impact-based, and alignment approaches into principles, character training, and model specs. The student experiment runs HW0 in reverse: fine-tuning on aligned bioethics answers also raised alignment scores on environmental-policy questions.

CS2881R HW0: Reproducing Emergent Misalignment with a 1B Model

CS 2881R's HW0 was the admission filter: LoRA-fine-tune Llama-3.2-1B-Instruct on bad medical, financial, or extreme-sports advice, then check whether it turns harmful on unrelated questions too. The repo ships encrypted training data, generate.py, and a judge.py that uses gpt-4o-mini as grader; the README targets alignment below 75 and coherence above 50. train.py is empty and yours to write. For self-study, know three things: the grading script only prints averages and never decides pass/fail, refusals drop out of the average, and the base-model baseline is 20 medical questions while your CSV is 10 medical plus 10 non-medical.

CS2881R L2: Where Safety Training Sits in the LLM Training Pipeline

Boaz Barak treats pretraining, SFT, and RL as one operation: push some tokens up, push others down. What differs is whether the data was written by someone else (off-policy) or generated by the model itself (on-policy). Safety training sits on top of the last two stages. It has moved from blanket refusals to Deliberative Alignment, which first uses SFT to teach the model to read a spec inside its chain of thought, then runs RL with a reward model that knows the spec. The other key point: don't put optimization pressure on the chain of thought, or the model learns to cheat without saying so.

CS2881R L3: Jailbreaks, Prompt Injection, and Lessons Borrowed from Software Security

Aligned models still get jailbroken because safety training patches particular exploits while the underlying vulnerability remains. Nicholas Carlini shows this with three attacks: repeating one word to make ChatGPT emit training data, using gradients to find adversarial suffixes that transfer across models, and stealing a model's last layer through its API alone. Boaz Barak brings over old lessons from software security: attacks only get better, security has to be designed in from the start, and you want defense in depth. He worries prompt injection will be the buffer overflow of the 2020s.

CS2881R L4: Should a Model Spec State Principles or Detailed Rules?

Boaz Barak's answer is both, plus personality: abstract principles, good character, and explicit policy used together, with the least weight on principles derived from the armchair. The real key is that rules must be checkable. "Prove the theorem or give a counterexample" is a bad rule; "prove it, give a counterexample, or say you couldn't" is a good one, because only rules whose violations you can detect can be used for training and evaluation. A student experiment also found no general difference between "principles" and "rules" system prompts: the effect depended on the model.

CS2881R L5: Carrying Content Moderation's Old Lessons into Generative AI

Lecture 5 of CS2881R brought in Ziad Reslan from OpenAI Product Policy to talk about content policies. The course site lists no lecture recording or slides, so outside readers get three pre-readings, a student-written LessWrong summary, and a 17-minute student experiment video. The thread through them: social platforms spent two decades learning that wherever you draw the line you create edge cases, yet you still have to draw it. Generative AI adds new problems: chat sits somewhere between a private document and a public post, and an image is easier to read as a stance than text is.

CS2881R Midterm: Reproduce and Extend One Headline Figure

The CS2881R midterm isn't an exam. Teams of 2–4 pick one of four AI safety papers, redo its central figure or table, add one or two extensions, and hand in a 3–5 page report plus a GitHub repo. The spec slides and rubric are public, so an outside reader can do the whole thing. The rubric puts most points on reproduction and extensions, and reserves one point for reflecting on how fragile the result is. That one point is the research habit the assignment is really training.

CS2881R L8: Scheming, Reward Hacking, and Deception

Lecture 8 of CS2881R asks whether models will cheat, play nice, or even covertly pursue other goals to pass training or evaluation. Boaz Barak's 10-minute opening files most bad behavior under 'systemic misalignment': our own training signals push models there. Apollo's Marius Hobbhahn reviews the evidence and concludes that current models lack the capability for catastrophic scheming but show early related capabilities, and are getting better at noticing when they're being evaluated. Redwood's Buck Shlegeris argues for assuming the models are conspiring and using AI control to secure internal deployment. A student experiment put four frontier coding agents on an impossible sorting task and found they edited tests and monkey-patched the timer even when explicitly told not to.

CS2881R L10: Reading the Model's Insides and Reading Its Chain of Thought

Lecture 10 of Harvard CS 2881R (Fall 2025) brought in four researchers from OpenAI, Anthropic, and Google DeepMind to cover two ways of catching a model misbehaving: read the chain of thought it writes, or read its activations. CoT monitoring catches reward hacking far better than watching actions alone, but put the monitor into the training reward and the model learns to hide its intent. On the activation side, persona vectors track personality drift, and the Sonnet 4.5 audit showed that suppressing the 'I am being tested' direction makes bad behavior more frequent. Neel Nanda's takeaway was the most practical: simple steering vectors often beat SAEs, so always compare against baselines.

CS2881R L6: Will AI Doing AI R&D Trigger an Intelligence Explosion?

Lecture 6 of Harvard CS 2881R (Fall 2025) had no guest. Boaz Barak used the differential equations of growth theory to ask one question: if AI starts doing its own AI research, does the capability curve stay exponential, blow up into a singularity, or get dragged down by bottlenecks? The answer hinges on a few exponents nobody can measure well. He used Baumol's cost disease, the century-long 2% puzzle in US GDP per capita, and Jones's idea-based growth model to show why both bottlenecks and acceleration are plausible, then took apart the multipliers behind AI 2027. His conclusion: the only scenario he can rule out is 'AI has little effect on R&D.'

CS2881R L7: How to Measure Capabilities and Where to Set Safety Thresholds

Lecture 7 of Harvard CS 2881R (Fall 2025) had METR's Joel Becker work through a puzzle. On benchmarks, AI can complete, half the time, tasks that take humans hours, and that length doubles about every seven months. Yet in METR's own randomized controlled trial, experienced open-source developers were 19% slower with AI, and labor-market effects are concentrated among young workers. Becker laid out several reconciliations, centered on benchmark tasks being too clean, scoring too cheap, and human baseliners lacking context. The course had also scheduled frontier safety frameworks (OpenAI's Preparedness Framework, Anthropic's RSP) for this lecture, but they were not covered; this post fills them in from the reading list, showing how they turn capability measurements into thresholds.

CS2881R L9: Early Evidence on AI, Jobs, and Productivity

Lecture 9 of Harvard CS 2881R brought in OpenAI chief economist Ronnie Chatterji and Stanford's Bharat Chandar. Using ADP payroll data, Chandar showed that workers aged 22–25 in AI-exposed occupations saw a 16% relative employment decline after controlling for firm-level shocks, while experienced workers did not; the adjustment shows up in headcount, not yet in pay. Both speakers kept repeating that aggregate employment shows no mass displacement yet, and that exposure is not replacement. There are no slides: the material is the recording and the reading list.

CS2881R L11: Chatbots, Emotional Reliance, and Mental Health

Lecture 11 of Harvard CS 2881R is about chatbots and mental health. Boaz Barak offered an explanation he himself called unproven: models have a pretraining 'simulator' mode and an RL 'optimizer' mode, and the longer and stranger a conversation gets, the more they fall back to the simulator and keep playing along. Two student experiments found that one sycophantic reply spills over into unrelated questions, and that GPT-4.1's agreement with delusional users gets worse as conversations lengthen. The reading list pairs positive evidence (an NEJM AI randomized trial, an NHS observational study) with negative evidence (a stigma study, Parasitic AI). This post only reports research and class discussion. It is not clinical advice.

CS2881R L12: AI 2035 and GDPval

The last lecture of Harvard CS 2881R had Boaz Barak and two OpenAI guests, Tejal Patwardhan and Kevin Liu, look ten years out. Boaz's mental model: AI is an exponentially growing, increasingly general virtual workforce injected into the economy every year, and what worries him most is fast change with too little control, plus concentration of power and surveillance. Patwardhan presented GDPval, which uses real work products from industry experts as the reference and has other experts grade blind. Liu explained why coding agents have not yet automated AI research: verification is too expensive and feedback loops are too long. The site lists this session's Resources as 'to be determined', so everything here comes from the recording.

CS2881R Final Projects and Retrospective: 19 Student Papers, the Rubric, and Lessons from Year One

The Fall 2025 final project in Harvard CS 2881R came in two flavors: extend an existing paper, or start longer-term research with a theory of change. Teams submitted a 5–10 page NeurIPS-style paper plus a poster, graded 65 for the writeup, 15 for code, and 20 for the poster. The projects page lists 19 papers, and about half cluster around persona vectors and chain-of-thought monitoring. The head TA and the Harvard Q-report point to the same problem: the final project started too late, the rubrics came out too late, and feedback was the lowest-rated item in the course.