Skip to content

CS2881R L12: AI 2035 and GDPval

Sep 30, 20261 min
TL;DRThe last lecture of Harvard CS 2881R had Boaz Barak and two OpenAI guests, Tejal Patwardhan and Kevin Liu, look ten years out. Boaz's mental model: AI is an exponentially growing, increasingly general virtual workforce injected into the economy every year, and what worries him most is fast change with too little control, plus concentration of power and surveillance. Patwardhan presented GDPval, which uses real work products from industry experts as the reference and has other experts grade blind. Liu explained why coding agents have not yet automated AI research: verification is too expensive and feedback loops are too long. The site lists this session's Resources as 'to be determined', so everything here comes from the recording.

🌏 中文版

Version note: This post is based on the November 20 session on the Harvard CS 2881R AI Safety Fall 2025 site and the Lecture 12 recording (YouTube title "Lecture 12: AI in 2035 and GDPval", about 1 h 47 min, uploaded January 2026). I checked every fact against the official materials on 2026-09-30. Recording content comes from YouTube's auto-generated captions. Materials for this lecture: the site has only the guest list, the recording link, and one line, "Discussion of future directions in AI safety research". Resources reads "Resources to be determined", and there are no slides. For papers and blogs the speakers mentioned in class, I add a link only when I am sure which one they meant. After the class break the recording cuts straight to Kevin Liu's segment, so the student experiment is not included. The series overview covers access grading for the whole course.

Series: previous L11: Chatbots, Emotional Reliance, and Mental Health | next Final projects and course retrospective | Series overview

L9 looked at traces AI has already left on the labor market, and L11 at an unexpected failure mode. The last lecture stretches the timeline ten years out. All three speakers work at OpenAI. Introducing the guests, Boaz said he is not such a fanboy that he thinks OpenAI does everything best, but he does think it has the best frontier evaluation team anywhere. Keep that in mind while reading; the course homepage also states Boaz's conflict of interest.

How the lecture is structured

SegmentSpeakerRecording time
The future of AI and his worriesBoaz Barak0:00–0:20
GDPval: measuring real workTejal Patwardhan (Kevin Liu presents part of the results)0:24–1:20
Why research is not automated yet; AI for scienceKevin Liu1:20–1:42
Open problems in pre-launch evaluationTejal Patwardhan1:42–1:45

Boaz: an exponentially growing virtual workforce

Boaz said his opening draws on his own blog post about AI and the economy (the Thoughts by a Non-Economist post covered in L9). His mental model fits in one sentence: think of AI as injecting an exponentially growing number of virtual workers into the economy every year, each year more capable and more general. Robotics may lag at first but will catch up.

He started from a figure in METR's task-horizon study (listed in the site's L7 reading list): tasks that are easy for humans eventually get solved nearly 100% of the time. He added a big caveat. That holds only in the non-adversarial setting. Adversarial robustness is unsolved, and you can always find inputs on easy tasks that make models fail.

His inference is that within about twenty years, maybe less, the global workforce effectively grows tenfold. The last time the global population grew tenfold, you have to go back to 1750. After the industrial and scientific revolutions, global GDP, life expectancy, and child mortality all improved dramatically, including in sub-Saharan Africa, which follows the same trend with a lag. His conclusion: if AI really multiplies the workforce by ten, the overall change is probably a radical one for the better.

What worries him

He said he is less worried about the classic "AI takes over and kills everyone" scenario. In his view we are neither in full control of models nor entirely in the dark but somewhere in the middle, still far from where he would feel comfortable. What really worries him is three things:

  • Fast change, too little control. We may squeeze 100 to 200 years of progress into 10 to 20, and society, government, and international relations are not ready. His analogy was aviation: it took 50 years to drive down accident rates, and planes barely changed in that time. With AI we are improving safety while going from a bicycle to a plane.
  • Concentration of power. AI can empower individuals, for example so a small nonprofit cannot be buried in paperwork, or government becomes more transparent. It can also enable surveillance at an unprecedented scale. Totalitarian governments were once limited by not being able to afford a spy for every citizen; now they might afford ten. If a government has perfectly obedient AI workers, none will refuse an unconstitutional order and none will blow the whistle.
  • Our private conversations with AI could be subpoenaed and surveilled unless the law changes.

He also said in passing that AI's environmental impact is real but often grossly exaggerated, and that job displacement is an issue but economic growth will help a lot: many US budget fights would vanish at 4% growth.

Patwardhan: how GDPval measures real work

Tejal Patwardhan helped build GDPval. She said earlier evals like MMLU and GPQA measured exam-style questions, and GDPval was one of the first ways to measure whether models can do real work.

Why measure capability instead of waiting for adoption data? Usage, productivity, and GDP are lagging indicators. Airplanes, railroads, and the internet took years or decades to go from existing to widespread; self-driving cars too, since it takes time and cultural change for people to get in one. Evals can forecast what models can do before regulation and culture catch up.

Where the tasks come from:

  • Start with sectors that contribute more than 5% of US GDP
  • Within them, pick the highest-earning occupations that are mainly digital work, such as editors, registered nurses, real estate brokers, and pharmacists
  • Recruit experienced practitioners (who pass resume screening, background checks, and more) to contribute work products they actually produced on the job, scrubbed of personal information
  • Tasks range from hours to weeks, e.g. a manufacturing engineer designing a cable reel stand, or a banking analyst building a competitor valuation landscape

How grading works: a separate panel of experts plays the manager, compares two deliverables blind (the human's real work and the model's output), and picks which they would rather use, weighing both correctness and subjective aspects like formatting and readability. Each task gets at least three experts, each looking at at least three model samples, and the win rate is an average, not best-of-3.

Results she reported in class:

  • The y-axis is the share of comparisons where the model wins or ties against an industry professional. GPT-4o was at roughly 10%; o3 and GPT-5 are clearly approaching expert level, and the trend looks roughly linear
  • In their evaluation Claude came closer to parity with experts
  • OpenAI models were stronger on pure-text tasks; Claude was stronger on tasks with multimodal files
  • Higher reasoning effort improves results
  • Simulating a "let the model try first, redo it yourself if unhappy" workflow: every model since GPT-4 saves time and money, with GPT-5 about 1.4–1.6x faster and cheaper

She spent time on how bad the failures are. Among samples where GPT-5 lost, experts re-rated them: about 23% actually thought the model was better (reflecting that experts disagree with each other), about half judged it acceptable but worse than the human, about 27% clearly subpar, and about 3% catastrophic. As long as the catastrophic rate stays there, she said, it is hard to deploy models without human supervision.

Where models go wrong: presenting this part, Kevin Liu said many failures come down to instruction following and formatting, not intelligence: not producing the requested file type, slide content running off the page, special characters that some fonts cannot render. They wrote a detailed prompt listing every common problem and added best-of-4, and scores rose meaningfully. His read is that there is a lot of low-hanging fruit.

Limits they listed themselves:

  • Digital work only, with a limited set of occupations
  • Tasks are self-contained and one-shot; it does not measure the context someone gains after two years at a job, where one sentence is enough to know what to do
  • Task prompts are very detailed because the model knows nothing about the job; Patwardhan said the paper has an experiment with less context, and performance drops but not by much
  • It measures a well-specified subset of tasks, not whole jobs

A public subset and an automated grading service are available, and Kevin Liu noted that OpenAI's other evals are listed at evals.openai.com. One student complained that a submission had gone unanswered for weeks. Patwardhan admitted the automated grader "is not that great" and said they are considering open-sourcing more of the eval so researchers can grade themselves.

"Are you scared?"

Patwardhan asked the class whether they felt scared. A few answers are worth keeping. One student who had left an entry-level finance job to return to school said it was precisely because tasks like GDPval's would catch up to that job. Another worried that not everyone can adapt, and that fear itself will hold people back. Another worried about concentrated power and asked whether frontier models might be held back to manage the disruption. Patwardhan answered that labor has more power when it is more valuable, and GDPval's tasks were designed from the Bureau of Labor Statistics' task breakdowns for each occupation, so if models keep improving on them, labor's leverage may become less unique.

Asked whether AGI is inevitable, Boaz said yes unless there is a surprising technical roadblock, citing the bitter lesson: the best coding models are trained on huge amounts of data unrelated to code. Liu thought timelines could be delayed, but avoiding it altogether would mean saying no every time across a million small "build a narrow system or hand it to the general one" decisions, which is very hard given how well society coordinates.

Liu: why coding agents have not automated AI research

Kevin Liu first showed SWE-style benchmarks, Terminal-Bench, and OpenAI's PaperBench (reproducing ICML papers) all rising steadily, then asked: so why hasn't research been automated? His reasons:

  • Research is often bottlenecked on compute and waiting for experiments, not on writing code
  • ML code gives slow feedback. In software engineering you rebuild and know in two minutes; an ML bug may take a day of running to confirm it is fixed
  • Verification is very hard. He gave two examples: a TPU-related bug Anthropic disclosed, where fixing one bug introduced a compiler bug that only appeared for certain batch sizes and model configurations; and OpenAI's investigation of whether Codex had gotten worse, which took regression analysis across hardware types to find. A degraded model still "works", just slightly worse, which is extremely hard to notice
  • You need to reproduce ablations from long ago, which means keeping every version of a repository giving the same results forever
  • Large-scale training infrastructure is complex. If each attempt takes two days, the gap between a model that needs 20 tries and a human who needs 5 really hurts

He observed that some very productive people at the company barely touch agents and instead write code right the first time in Vim, because verification is so expensive.

Where models already help a lot, he summarized, the work tends to be hard to do but fast to verify, tolerant of failure, fuzzy in its correctness, light on deep context, and heavy on persistence rather than raw brainpower. His examples were agents that audit other models for misbehavior (he mentioned related work from Transluce and Anthropic) and Codex code review: purely additive, cheap to check, and occasionally catching something interesting.

His predictions:

  • "Vibe-coded" research code (data processing, plotting, sweeps) will be the majority by volume and written by models, but that does not mean 80% of the important work speeds up
  • Today's world is built for "two-second agents" (deterministic programs) and "two-year agents" (engineers who stay two years); "two-day agents" in between are hard to use. Either people put a lot of manual work into adjusting workflows, or models' time horizons get longer
  • Predictable scaling in capability does not mean predictable impact. Real tasks chain many subtasks, and one missing piece blocks the whole thing, so real impact tends to arrive in steps

On AI for science, he said models are currently good at literature review, turning over every rock, and well-defined math of moderate difficulty. Boaz added his own usage: having Codex handle LaTeX and bibliography, and read through a whole paper flagging unclear spots, though not yet for original research. He also warned that handing the "obvious" steps to a model cuts both ways, since re-walking trivial derivations is sometimes part of learning; that is why he tells new team members to ask people first.

The last three minutes: three open problems in pre-launch evaluation

Patwardhan closed with two slides and urged students to go read any model's system card. She listed three things that get harder as capability grows:

  1. Capability elicitation: with more test-time compute, better scaffolds, or better prompts, models can do more. How do you make sure pre-launch testing actually reaches the ceiling?
  2. Situational awareness: models may realize they are being tested, so test results may not reflect behavior in deployment
  3. Long-horizon risks are hard to test before launch: a bioattack plan spanning months of wet-lab work cannot be measured in a time-boxed eval. We need scaling laws that extrapolate long-horizon capability from short-horizon evals

Her closing point was that many of these research questions still have no answers.

What to watch for in this lecture

  • All three speakers work at OpenAI. The GDPval results, the failure breakdown, and the "models keep getting better" judgment are the developer's oral report in class. Numbers here come from auto-captions, so check the original paper before citing them
  • The site has no reading list, so the book review Boaz mentioned and the blogs Liu cited are not linked here
  • The takeaway is less the predictions themselves than three ways of looking ahead: Boaz through macroeconomic analogy, Patwardhan through task evaluation, Liu through bottlenecks in research workflows

How to study it

  1. Read the Boaz blog post section in L9 first, then watch Boaz's opening. His "virtual workforce" is the spoken version of that post.
  2. While watching Patwardhan, compare with the L9 student experiment. The students used a model as grader; GDPval uses blind expert grading. What changes?
  3. While watching Liu, apply his list of "traits of work that suits models" to one task on your own plate.

One thing to do tonight: following Patwardhan's advice, open the system card of a model you use and read only the table of contents. Assign each section to one of her three problems. The sections that fit none of them show where you do not yet know what the evaluation team worries about.

Further reading

References