CS231N L9: Object Detection, Image Segmentation, and Visualization
Lecture 9 of CS231N Spring 2026 moves from one label per image to one label per pixel and per object. Semantic segmentation uses fully convolutional networks that downsample and then upsample, and U-Net feeds high-resolution features back in. Detection goes from R-CNN's roughly 2,000 CNN forward passes to Fast R-CNN, Faster R-CNN's RPN, single-stage YOLO, and anchor-free DETR. Mask R-CNN adds a 28×28 mask per RoI. The last part covers saliency, CAM, and Grad-CAM. The adversarial examples, DeepDream, and style transfer listed on the schedule appear in neither the 2026 nor the 2025 slides.