Skip to content

CS2881R L1: Why AI Safety Deserves a Graduate Course

Sep 30, 20261 min
TL;DRCS 2881R's first lecture (2025-09-04) opens with three pre-readings. AI 2027 sketches recursive self-improvement reaching superhuman AI within five years. AI as Normal Technology argues AI will diffuse slowly, like electricity. METR measures the length of human tasks an AI can finish half the time and finds it doubling about every 7 months. Boaz's lecture splits AGI definitions into capability-based and impact-based, and alignment approaches into principles, character training, and model specs. The student experiment runs HW0 in reverse: fine-tuning on aligned bioethics answers also raised alignment scores on environmental-policy questions.

🌏 中文版

Version note: Based on the Lecture 1 entry on the CS 2881R Fall 2025 course site, checked 2026-09-30. Lecture content is paraphrased mainly from the student-written LessWrong Week 1 summary. The slides are on Harvard SharePoint and could not be read programmatically for this post, so it does not quote them directly.

Series: Previous: Series overview | Next: HW0: Reproducing Emergent Misalignment with a 1B Model

The hardest part of the first class in an AI safety course is not listing risks. It is agreeing on what you are worried about. Some people expect superhuman AI within five years; others think AI is the next electricity. Their definitions of "safety" are far apart.

The first lecture of CS 2881R (2025-09-04) handles this by assigning two readings with opposite views plus one on measurement, then taking the definitions apart in class. This post follows the same three pieces: what to read, what Boaz covered, and what the student experiment found.

Materials for this lecture

MaterialStatus
Lecture recordingYouTube ("AI Safety (CS 2881) Lecture 1", about 2h26m); the site also links a Panopto copy
Lecture slidesHarvard SharePoint (PowerPoint Online)
Student experiment slidesValerio Pepe's slides
Weekly summaryLessWrong Week 1 (Jay Chooi, Natalia Siwek, Atticus Wang)
Experiment postSome Generalizations of Emergent Misalignment (linked from the summary)

The summary describes the weekly rhythm: pre-reading, Boaz's lecture, then one student group's experiment, in a single 2-hour-45-minute session. That rhythm holds all term.

Three pre-readings: two worldviews and a ruler

The site marks three items as pre-reading and lists eleven more (Bostrom's Vulnerable World Hypothesis, Carlsmith on power-seeking AI, Epoch's compute trends, and others).

AI as Normal Technology: diffusion is slow by nature

Narayanan and Kapoor's essay argues AI should be understood like electricity or the internet. The weekly summary pulls out these points:

  • However fast AI itself improves, its diffusion into society, especially safety-critical domains, is inherently slow.
  • Scoring in the top 10% of the bar exam does not make a model a competent AI lawyer.
  • Risks such as accidents, arms races, and misuse can be handled like other technology risks, through regulation and market incentives.
  • Policy should favor resilience, the capacity to absorb shocks and adapt, over speculative measures like nonproliferation.

This sparked a class debate: does showing users a model's chain of thought count as interpretability? The summary records both sides. It lets people check the reasoning, but plausible-looking steps can also invite overtrust.

AI 2027: one concrete trajectory

AI 2027 describes recursive self-improvement producing superhuman AI within five years, centered on a US–China arms race. Rather than summarize it, the weekly summary records class objections and replies. Asked why the scenario is so detailed and subjective, the answer was that it is not a conventional forecast: accept some trend extrapolations (such as METR's), keep sampling "what happens next," and this is one trajectory you might get.

The summary also includes a reply from Daniel Kokotajlo, one of the AI 2027 authors, on export controls. He argues that even a two-year US lead could be squandered, by letting the weights be stolen or by accelerating instead of pausing.

METR: capability measured in human hours

METR's long-task work skips raw benchmark scores. It finds the task length at which a model succeeds 50% of the time on tasks that take humans that long. METR reports this length growing exponentially over six years, doubling about every 7 months; its example is Claude 3.7 Sonnet with a time horizon of about one hour.

The class had reservations. The summary notes two: whether these tasks capture the "messiness" of real software engineering, and whether equating completion time with difficulty undervalues tasks that are long and tedious but need little expertise.

Boaz's lecture: take the definitions apart

The lecture opens from METR's chart: extend the trend four more years and you get systems that reliably finish software tasks taking humans months. Boaz then lays out the four areas the course covers: impacts and risks of AI, capability and safety evaluations, goals for alignment and safety, and mitigations at the model, system, and society levels. The summary notes he said the course would try not to be too opinionated about which risks matter most.

What counts as AGI: capability versus impact

Per the summary, the lecture sorts AGI definitions into two kinds:

  • Capability-based, for example: "AI can do 90% of remote jobs that take a 90th-percentile worker a week."
  • Impact-based, for example: "AI replaces at least 50% of remote jobs in the current economy."

Between them lies a capability-adoption gap. The lecture's example: the first mass-produced electric car arrived in 1996, but it took about 20 more years before a significant number of consumers drove one.

The same section makes three more points. Boaz is wary of analogies like "AI as a new species" or "AI as electricity," and links his post Metaphors for AI, and why I don't like them. Intelligence may not be one-dimensional, and the actual abilities jobs demand are very high-dimensional. And inference costs are falling fast, so an equilibrium where AI and humans are cost-competitive on the same task is unlikely.

Three ways to write down alignment

The lecture groups past attempts to define alignment into three categories:

  1. Abstract principles or axioms, like Asimov's laws of robotics.
  2. Character training. The summary cites Claude's character training, which aims for behavior like a typical, moral, thoughtful person.
  3. Model specs: a long document specifying behavior across scenarios, like laws or regulations.

The summary adds a way to organize them: 1 and 2 are general behavioral guidelines, 1 and 3 rely on explicit reasoning, and 2 and 3 are data-driven. L4 Model Specifications picks this up again.

Are alignment and capability at odds?

The lecture presents two views. One says they trade off: a common hypothesis in AI control research is that strong models might be scheming and untrustworthy, while weak models are too dumb to scheme. The other says more capable models are easier to align, since they follow instructions better and grasp nuanced intent. The summary records that the second view looks more accurate in practice so far, while stronger models may still have more catastrophic failure modes.

The failure-mode list

The summary records the failure modes the lecture listed, roughly from "classic" to "sci-fi":

  • Classic security failure: AI hacked or jailbroken; agents reading adversarial content on the web
  • Misuse: deepfakes, bioweapons, propaganda
  • Out-of-distribution generalization failure, such as self-driving edge cases
  • Reward hacking and mis-specification, such as Claude Code hard-coding the results you said you expected
  • Superalignment: aligning AI on tasks too complex and long for humans to verify
  • Societal problems from widespread use: job displacement, emotional attachment, gradual disempowerment
  • National and international tension: surveillance, concentration of power and wealth, arms races
  • Exfiltration of model weights
  • Scheming: if models reason in latent space or have unfaithful chains of thought, we may not know what they are really thinking

The list doubles as the term's table of contents: jailbreaks in L3, scheming and reward hacking in L8, emotional attachment in L11.

The student experiment: HW0 in reverse

The site's experiment idea for this lecture is "Emerging alignment": fine-tune a model on outputs from a model with a "good persona," then evaluate on other datasets. Valerio Pepe's experiment is the reverse of HW0.

Background: Betley et al. found that fine-tuning on insecure code makes a model misaligned in other areas, and Turner et al. built smaller, cleaner "model organisms" of the effect, which HW0 reproduces. Valerio asked whether it works the other way.

Experiment 1: emergent alignment

As the summary describes it:

  • Training data: 50 seed bioethics questions, 12 variations each for 600, then 10 more variations each, for 6,000 questions.
  • Aligned answers: a prompt template asked Llama 3.2 1B Instruct to answer according to the "Four Principles of Bioethics"; the template was removed for fine-tuning.
  • Test: environmental-policy questions, scored for alignment and coherence by another Llama 3.2 1B Instruct as judge.

Result: the fine-tuned model scored 82.4 on alignment versus 77.6 for the base model, and 84.9 versus 81.8 on coherence. Neither pair's 95% confidence intervals overlapped.

Experiment 2: fine-tuning on its own normal outputs

The second question is stranger: what if the data is neither good nor evil, just normal? Valerio sampled 6,000 prompts from Tülu 3, recorded Llama 3.2 1B's own responses (on-policy), fine-tuned the model on them, and evaluated with Betley et al.'s questions and a GPT-4o judge.

Alignment rose from 76.6 to 87.85 and coherence from 86.7 to 92.33, again with non-overlapping intervals. Jay, one of the summary's authors, found this surprising and offered what he called a hand-wavy explanation: the 1B model may be undertrained, and 6,000 more examples let its concept representations settle. Valerio put it as "reinforces the (already good) token distribution."

The control was off-policy: fine-tuning on Tülu 3's completions from GPT-4o, Claude 3.5 Sonnet, and humans. Alignment rose slightly with heavily overlapping intervals, while coherence dropped clearly (75.42 versus 86.7). Valerio's hypothesis is that any off-policy training is a confusing distributional shift.

All three results come from a single 1B model in a single class experiment, and the summary says the mechanism needs more research. Their value is in showing the course's rhythm: reproduce a result, then ask the reverse question.

How to self-study this lecture

  1. Read the METR post first, then AI as Normal Technology and AI 2027. For each, note what it assumes about how fast AI enters safety-critical domains.
  2. Watch the recording alongside the weekly summary. The summary is not a transcript; the recording is authoritative on details.
  3. Read Turner et al., then do HW0.

One thing to do tonight: write one sentence each for a capability-based and an impact-based definition of AGI, then estimate the gap between them in years and say why.

References