Skip to content

CS231N Wrap-Up: World Modeling / Robot Learning, Human-Centered AI and the Final Project

Sep 30, 20261 min
TL;DRThe last two lectures of CS231N Spring 2026 have no public slides. The schedule lists L17 only as "World Modeling" with guest lecturer Gordon Wetzstein, and L18 only as "Human-Centered AI." Outside readers get 2025 substitutes: that year's L17 was a different topic, Robot Learning (Yunzhu Li, slides and video), and L18 is a Fei-Fei Li recording with no slides. This post labels each year separately and never presents 2025 content as 2026. The second half covers the final project: 35% of the grade, two tracks (Applications and Models), pixels required, and deliverables of a one-paragraph proposal, three milestone check-ins, a 6–8 page report and a poster.

🌏 中文版

Which year: The years in this post are the messiest in the series, so here they are up front.

  • 2026 L17 "World Modeling" and L18 "Human-Centered AI": only entries on the Spring 2026 schedule, with no slide links. On 2026-09-30, both slides/2026/lecture_17.pdf and lecture_18.pdf returned 404. The 2026 recordings are on Canvas only.
  • 2025 L17 "Robot Learning": slides (103 pages) and a recording (about 1h18m), lecturer Yunzhu Li. A different topic from 2026.
  • 2025 L18 "Human-Centered AI": a recording only (about 1h05m), lecturer Fei-Fei Li; the slide URL also returns 404. The L18 summary below comes from the video's English captions.
  • Final project: the Spring 2026 project page and section 3 slides.

This is post 20, the last one, in the Reading Stanford CS231N series.

Series: previous L15: 3D Vision | series overview

The last two lectures stop teaching new algorithms. They push what came before outward: how vision models connect to action, and whom they should serve. Meanwhile, the heaviest thing on enrolled students' plates is the final project, worth 35% of the grade.

2026 L17: World Modeling, one line on the schedule

For May 28, 2026, the schedule has two lines: "Lecture 17: World Modeling" and "Guest Lecturer: Prof. Gordon Wetzstein." No subtopics, no slides, no suggested readings. A3 is due the same day.

That is everything an outside reader can confirm. This post doesn't guess what the lecture covered. The 2025 Robot Learning lecture below is a different topic with a different speaker. Treat it only as a reference for how the course once ended at this point.

2025 L17: Robot Learning (Yunzhu Li)

In 2025 this slot went to guest lecturer Yunzhu Li from Columbia. The schedule lists Deep Reinforcement Learning, Model Learning and Robotic Manipulation. The slides have seven parts: problem formulation, robot perception, reinforcement learning, model learning and model-based planning, imitation learning, robotic foundation models, and remaining challenges.

From supervised learning to agents that act

The opening places the whole course in one frame. So far you've seen supervised learning (x and y) and self-supervised learning (x only). Now an agent takes actions in an environment and receives rewards, and the goal is to learn actions that maximize reward. Examples run from cart-pole, robot locomotion, Atari and Go to text generation, chatbots, and a cloth-folding robot.

The slides then explain how robot vision differs from the computer vision in earlier lectures: it is embodied, active and situated. The robot has a body, so its actions feed back into its own sensing immediately. It knows why it wants to look and chooses what to perceive. The lecture's key challenge fits in one line: close the perception–action loop.

Why RL differs from supervised learning

The slides give four reasons. Rewards and state transitions may be random. The reward may not depend directly on the current action (credit assignment). The world isn't differentiable, so you can't backprop through it. And what the agent sees depends on how it acts (nonstationarity). Case studies run from DQN on Atari and the AlphaGo line to quadruped locomotion and a robot hand solving a Rubik's Cube.

Model-free RL's problems are listed plainly: trial and error, lots of interaction, safety concerns, little interpretability. The response is model learning and model-based planning. Learn a dynamics model of the world, plan through it, execute the first action, observe the new state, and re-optimize. The key question is what form the state should take. The slides compare pixel, keypoint and particle dynamics, including RoboCook, which manipulates elasto-plastic objects with diverse tools using particle dynamics.

Imitation learning and robotic foundation models

Imitation learning is supervised learning from demonstrations. The slides list behavior cloning, iterative collection of expert demonstrations, inverse RL, implicit behavior cloning, and diffusion policies.

Next come robotic foundation models: a policy that maps (observation, goal) straight to action, with no explicit states or transition functions, also called VLAs or large behavior models. The slides give a timeline from RT-1 (December 2022) through Pi-Zero, OpenVLA, Gemini Robotics, GR00T and others, and use Physical Intelligence's Pi-Zero to show pre-training and post-training on cross-embodiment data.

Pointing toward world models

Two points from the challenges section are worth keeping. First, evaluation: it happens mostly in the real world, which is costly and noisy, and training loss correlates only weakly with real-world success; simulation has its own sim-to-real gap. Second, one slide titled "Foundation Policy → Foundation World Models," where the speaker defines a world model as action-conditioned future prediction, with examples like 1X World Models, DayDreamer and NVIDIA Cosmos.

The 2026 L17 happens to be titled World Modeling. That is a connection I see between the two schedules and the 2025 slides. It doesn't mean the 2026 lecture covered this material.

L18: Human-Centered AI (2025 recording only)

The 2026 L18, on June 2, is just a title on the schedule. In 2025, Fei-Fei Li gave this lecture. There are no public slides, but the recording is complete. What follows is based on its captions.

She opens by saying this lecture teaches no new algorithms. It's a talk about long-term research evolution and the human perspective, titled "What we see and what we value: AI with the human perspective." It has three parts:

  1. Building AI to see what humans see. From the origin of vision 540 million years ago and the 1960s summer vision project to three waves of object recognition: part-based models inspired by psychology, then statistical machine learning, then ImageNet, CNNs and GPUs converging in 2012. After that come relationships (scene graphs, Visual Genome), image captioning and dense captioning, and still-unsolved multi-actor activity understanding in video. Her conclusion: the field has always drawn on cognitive science and neuroscience, and it will keep doing so.
  2. Building AI to see what humans don't see. On one side, superhuman ability, such as fine-grained recognition of bird species and car models, and using Street View car models to study social patterns. On the other, human limits, such as limited attention leading to medical errors, with AI counting gauze during surgery (she stresses this is a demo, not a deployed system). Then bias: human vision is biased, so is data, and AI can amplify both. And privacy: some things shouldn't be seen, illustrated by a hardware–software approach that recognizes actions while protecting privacy.
  3. Building AI to see what humans want to see. Starting from labor anxiety, she argues for augmenting rather than replacing people. Her examples are ambient intelligence in health: depth-only sensors that monitor hand hygiene, track ICU patient mobility, and help seniors age in place. Finally, robots: using LLMs and VLMs to plan actions from open-ended instructions, and the BEHAVIOR benchmark, which first asked about 1,400 people which household tasks they want robots to do and then built simulation environments around them. She says they tested three BEHAVIOR tasks with current robot algorithms, and without privileged information, performance was zero.

Her closing message: AI should be a tool that augments people, not one that replaces them.

The final project: 35% of the grade

The project page frames the project as applying what you learned to a problem you care about. The one hard rule is that it must involve pixels in some form. A pure NLP project doesn't qualify even if it uses ConvNets; related areas that vision conferences accept, like shape analysis, are allowed.

Two tracks

TrackOfficial description
ApplicationsBring your background (biology, engineering, physics) and apply the course's vision models to a real problem in your domain
ModelsBuild a new model or a variant of an existing one for vision tasks; harder, and sometimes publishable

Teams have up to 3 people, and solo work is allowed. The page says 3-person teams should deliver a more impressive write-up and results. It also warns that the baseline requirements for a complete paper have historically been hard for individuals without prior experience.

Deliverables and weights

DeliverableWeightDue (2026)Late days
Project Proposal1%Apr 23Yes
Milestone 1: Problem + Related Work3%May 15Yes
Milestone 2: Technical Approach3%May 22Yes
Milestone 3: Preliminary Results3%May 29Yes
Final Report20%Jun 5No
Poster session (in person) + poster PDF and code5%Poster Jun 10; PDF and code Jun 9No

Some details:

  • The proposal is one paragraph of 200–400 words: the problem, the reading, the data, the method, and how you'll evaluate.
  • The three milestone check-ins are new in 2026. Each is a 10-minute discussion in TA office hours. You submit 2–5 slides first, present for about 5 minutes, and take questions for the rest; all members attend. Each 3% splits into Progress, Clarity and Robustness at 1% each. Milestone 1 requires discussing at least 3 related papers.
  • The final report is 6–8 pages in the CVPR template. The rubric weights Introduction 10%, Related Work 10%, Data 10%, Methods 30%, Experiments 30%, Conclusion 5%, and Writing/Formatting 5%.
  • The report must cite any base code you used, including CS231N assignment code. A project shared with another class has to state which part counts for CS231N, and you can't submit the same PDF to both.

Generative AI policy

Using generative AI to produce project code follows the same rules as using public sources, and all use must be documented explicitly: plans, prompts, transcripts, and a marker on every AI-generated artifact. Using generative AI to write the final report violates the Honor Code. It may be used only for editing and formatting.

Section 3 slides: what counts as a good project

The section 3 slides (13 pages) add criteria the project page doesn't spell out. You don't need strict novelty or a state-of-the-art result. You do need real effort, and you should interpret results from several angles, not just plot a loss curve.

They describe weaker projects in two ways: spending weeks collecting and cleaning data without testing any hypothesis, or cloning a repo and stitching it together with no real contribution.

They sort recent successful projects into three patterns:

  1. Domain adaptation: apply a strong vision model or VLM to a meaningful new problem, with real task-specific data and evaluation. Examples include medical imaging, remote sensing, scientific imaging and sign language.
  2. Method improvement: start from a recognized baseline and make a technically meaningful change, such as a new loss, module or training method.
  3. Reproduction: rebuild a proprietary or hard-to-reproduce capability, where the implementation itself is the contribution.

The slides also advise on reading papers. Don't read linearly on the first pass: read the abstract word for word, then skim the figures and captions. If it's still relevant, read methods and results. Read the whole paper only when the detail genuinely helps.

How a self-learner can borrow this

Grading, TA check-ins and the poster session belong to enrolled students. But the structure works for self-study: past reports are public, and the proposal questions and report rubric are on the page.

  1. Pick a direction using one of the three section 3 patterns, and write a 200–400 word proposal that answers each of the five official questions.
  2. Set yourself three checkpoints, with 2–5 slides each, following the Milestone 1–3 requirements.
  3. Write the report against the rubric's seven sections. Methods and Experiments are 60% of it, so budget your time the same way.

One thing to do tonight: open the Spring 2025 report list, pick three reports, and decide which section 3 pattern each one fits.

End of the series

That completes this series: the public L1–L16 slides of CS231N Spring 2026, all three assignments, and L17 and L18, which survive only as schedule entries and 2025 recordings. The series overview has the full self-study plan and where each post fits.

Further reading

References