Skip to content
All tags

#embodied-ai

6 posts

CS224R L15: Hierarchical RL and Imitation Learning

Long-horizon tasks are hard because the agent visits a huge number of states and has many chances to make mistakes or get stuck. Lecture 15 of CS224R answers with two levels: a high-level policy proposes subgoals, and a low-level policy runs at a higher frequency to reach them. The real design decisions are three: how to represent the subgoal, how to supervise each level, and when to switch to the next subgoal. The slides also admit that nobody has yet shown whether hierarchy beats a single policy with chain of thought.

CS224R L17: RL for Robot Foundation Models (VLAs)

VLAs trained only with imitation learning often plateau around 80% success, while autonomous robots often need 99%+. Lecture 17 of CS224R splits "how do you improve a VLA with RL on a real robot" into three routes: recast RL as supervised learning (iterated offline RL), learn a small separate policy on the VLA's representation or diffusion noise, or learn a small policy that edits the VLA's actions. The slides call this an open research problem and describe the content as recent themes plus the speaker's opinion.

CS224R L16: Sim-to-Real Robot Learning

Simulators are cheap, fast, and safe, and they hand you labels the real world never will, but they never match reality exactly. In Lecture 16 of CS224R, CMU's Guanya Shi sorts the ways to close that gap into three families: domain randomization trains one policy that works across many physical parameters; teacher-student trains a teacher on privileged information and then has a student that sees only real sensors imitate it; real2sim2real uses real data to make the simulator more faithful. The advanced topics are defining tasks from human motion data and choosing RL algorithms suited to sim2real.

CS231N Wrap-Up: World Modeling / Robot Learning, Human-Centered AI and the Final Project

The last two lectures of CS231N Spring 2026 have no public slides. The schedule lists L17 only as "World Modeling" with guest lecturer Gordon Wetzstein, and L18 only as "Human-Centered AI." Outside readers get 2025 substitutes: that year's L17 was a different topic, Robot Learning (Yunzhu Li, slides and video), and L18 is a Fei-Fei Li recording with no slides. This post labels each year separately and never presents 2025 content as 2026. The second half covers the final project: 35% of the grade, two tracks (Applications and Models), pixels required, and deliverables of a one-paragraph proposal, three milestone check-ins, a 6–8 page report and a poster.

2025 AI Conference Review: Computer Vision

2025 was a two-conference year for computer vision, with CVPR and ICCV both taking place. CVPR received a record 13,008 submissions; Best Paper VGGT turned 3D reconstruction from iterative optimization into feed-forward inference. ICCV's Marr Prize went to BrickGPT, which generates brick structures from text that can actually be assembled. 3D Gaussian Splatting displaced NeRF, video generation moved toward products, and flow models began replacing diffusion, completing several paradigm shifts in one year.

aideep-dive

Where 3D Generative Models Stand: Reading the 2026 Technical Map Through Lyra 2.0

The dominant paradigm in 3D generation in 2026 is video diffusion feeding feed-forward 3D reconstruction, and Lyra 2.0 is the flagship of that line. But three Best Papers at CVPR 2026 point at what comes next: SAM 3D brings foundation-model-scale object reconstruction, D4RT rebuilds dynamic 4D scenes in seconds from a unified transformer, and O-Voxel replaces Gaussians with structured latents. 3DGS still rules, but surface primitives are challenging it, and pixel-space diffusion is pushing back against latent space.