A guided reading of Stanford CS231N (Deep Learning for Computer Vision), based on the Spring 2026 slides and assignments A1–A3, with the public Spring 2025 YouTube lectures for video. It runs from image classification, backprop, CNNs and Transformers through detection and segmentation, self-supervised learning, generative models, vision-language and 3D.
CS231N is Stanford's deep learning course for computer vision. For Spring 2026, slides for 16 lectures, all three assignment pages with starter code, the course notes, and the project spec are public, so this series rates it A3 (enough to self-study). There are three gaps: the 2026 recordings are Canvas-only, L17 and L18 have no slides, and the midterm is not public. You can pair the 2026 slides and assignments with the 2025 YouTube recordings and follow the official calendar over 10 weeks.
The first CS231N lecture of 2026 comes in two slide decks. The first tells the history of vision and deep learning on a single timeline: Hubel & Wiesel's cat experiments, Marr's stages of visual representation, then the Neocognitron, backprop, and LeNet, until ImageNet and AlexNet join the two threads. The second covers the course map, grading, and rules, and moves every assignment onto Colab. After this lecture you'll know which gap each of the remaining 17 lectures fills.
L2 starts from one question: a computer sees a grid of numbers between 0 and 255, so how does it recognize a cat? Hand-written rules don't scale, so the course switches to a data-driven approach: collect data, train, evaluate on new images. The first classifier, kNN, teaches how to split train/val/test, but pixel distances carry no meaning. The second, the linear classifier f(x,W)=Wx+b, can be read three ways (algebraic, visual as templates, geometric as hyperplanes). Softmax turns its scores into probabilities, and the loss is the negative log probability of the correct class.
L2 gave us a score function and a loss. L3 answers how to find a good W. The first half covers regularization: add λR(W) next to the data loss so the model does not fit the training data too well. The second half is a lineage of optimizers. SGD zigzags in narrow valleys, Momentum builds up velocity, RMSProp scales each dimension's step, Adam combines the two and adds bias correction, and AdamW moves weight decay outside the moment estimates. The slides close with practical advice: Adam(W) is a good default in many cases, and SGD+Momentum can do better but needs more tuning of the learning rate and schedule.
The first half of L4 replaces the linear classifier f = Wx with a two-layer network f = W₂ max(0, W₁x) and shows that dropping the max activation collapses it back into a linear classifier. The second half answers how to compute gradients once the network gets deep: draw the function as a computational graph, and each node only needs its own local gradient multiplied by the upstream gradient coming back from later nodes. Add distributes, mul swaps, max routes, copy sums, and those four patterns let you trace any network. The lecture ends with matrices: dL/dx always has the same shape as x, so never build the Jacobian.
CS231N Assignment 1 is worth 12% of the grade and was due April 16, 2026. All five Colab notebooks are hand-written numpy on CIFAR-10. Q1 kNN asks for distance computations with two loops, one loop, and no loops. Q2 Softmax goes from naive to vectorized to SGD. Q3 assembles affine, ReLU, and softmax into a two-layer network. Q4 switches to HOG and color-histogram features. Q5 generalizes to any depth and implements Momentum, RMSProp, and Adam. The 65 KB starter code is public; the Gradescope grading isn't. This post covers structure and goals only, not solutions.
A two-layer network flattens a 32×32×3 image into a 3072-dimensional vector, and the spatial structure is gone. CS231N Lecture 5 answers with two layers. A convolution layer slides small filters across the image and reuses the same weights at every position. A pooling layer downsamples and has no learnable parameters. Both are translation equivariant. One formula gives every layer's output size: (W − K + 2P) / S + 1.
The 2026 slides for CS231N Lecture 6 are titled "Training CNNs and CNN Architectures" and split into how to build and how to train. Only two architectures get case studies. VGG shows that three 3×3 convs are deeper and cheaper than one 7×7. ResNet lets layers learn the residual F(x) = H(x) − x, which fixes an optimization problem where deeper plain nets had worse training error. The most practical takeaway is transfer learning: with fewer than about a million images, start from a model pretrained on a large dataset.
An RNN updates one hidden state at every time step with the same weights, so it can handle sequences of any length. CS231N Lecture 7 starts by hand-building an RNN that detects repeated 1s. It then covers character-level language models and feeding CNN features into an RNN for image captioning. Gradient flow explains why vanilla RNNs are hard to train: clip gradients to stop them exploding, and change the architecture (the LSTM) to stop them vanishing. The lecture ends by calling state space models like Mamba "modern RNNs."
Assignment 2 of CS231N Spring 2026 is worth 18% of the grade. Across five notebooks you hand-write BatchNorm/LayerNorm, dropout, and the forward and backward passes for convolution and pooling, then learn PyTorch at three levels of abstraction, and finish with RNN image captioning on COCO in PyTorch. Q4 is the turning point: through Q3 you derive every gradient yourself, and from Q5 on autograd takes over while a numerical gradient check confirms it. The official slides warn that this is the longest of the three assignments.
Lecture 8 of CS231N Spring 2026 starts from the bottleneck in RNN translation models, abstracts attention into an operation on sets of vectors, builds up to self-attention, masking, and multiple heads, and shows the whole layer is four matrix multiplies. A Transformer block is self-attention, LayerNorm, residual connections, and an MLP; ViT turns a 224×224 image into 16×16 patches used as tokens. The lecture closes with four common post-2017 changes: Pre-Norm, QK-Norm, SwiGLU, and MoE.
Lecture 9 of CS231N Spring 2026 moves from one label per image to one label per pixel and per object. Semantic segmentation uses fully convolutional networks that downsample and then upsample, and U-Net feeds high-resolution features back in. Detection goes from R-CNN's roughly 2,000 CNN forward passes to Fast R-CNN, Faster R-CNN's RPN, single-stage YOLO, and anchor-free DETR. Mask R-CNN adds a 28×28 mask per RoI. The last part covers saliency, CAM, and Grad-CAM. The adversarial examples, DeepDream, and style transfer listed on the schedule appear in neither the 2026 nor the 2025 slides.
CS231N Lecture 10 treats video as 2D plus time, a T×3×H×W tensor, and follows one thread: efficiency. Train on short clips and average several clips at test time. Architectures run from per-frame 2D CNNs and late fusion to 3D CNNs, then two-stream networks that isolate motion with optical flow, and I3D, which inflates 2D weights into 3D. After 2021 the field moved to Transformers, where token counts explode, which led to divided space-time attention, Video Swin, MViT, and tubelets. The last part covers temporal localization, audio-visual models, VideoLLMs, and long-form video, where HourVideo shows how far the field still has to go.
CS231N Lecture 11 uses Llama3-405B as its running example. It starts with GPU hardware and clusters (the H100, 8-GPU servers, a 24,576-GPU cluster), then maps the four dimensions of a Transformer activation to four kinds of parallelism: split the batch for data parallelism (which grows into FSDP and HSDP), the sequence for context parallelism, the layers for pipeline parallelism, and the channels for tensor parallelism. Along the way it covers activation checkpointing (trading recomputation for memory) and a practical scaling recipe, and it uses Model FLOPs Utilization (MFU) as the tuning target: above 30% is good, above 40% is excellent.
CS231N Lecture 12 asks whether we can learn good representations without huge manually labeled datasets. The answer comes in three parts. First, pretext tasks that generate labels from image transformations: predicting rotation, solving jigsaw puzzles, inpainting, colorization, and MAE with a 75% mask ratio. Second, the more general contrastive learning: the InfoNCE loss, SimCLR with its large batches, MoCo, which decouples batch size from the number of negatives with a queue, and sequence-level CPC. Third, DINO, which needs no negatives: a student predicts the output of a momentum teacher, and centering plus sharpening prevent collapse. The core evaluation is linear probing: freeze the encoder and train only a linear classifier.
The first generative-models lecture of CS231N Spring 2026 starts by pinning down the difference between a discriminative model, which learns p(y|x), and a generative model, which learns p(x). Every possible image competes for the same probability mass, so a generative model can reject unreasonable inputs. A taxonomy then splits generative models into those that can compute p(x) and those that can only sample. The lecture covers the two that compute (or approximate) it: autoregressive models factor p(x) with the chain rule into step-by-step predictions, and their weak point on raw pixels is speed; VAEs cannot compute p(x), so they maximize a lower bound, the ELBO, whose reconstruction and prior terms pull against each other. The schedule lists GANs under this lecture, but both the 2026 and 2025 slides place them at the start of the next one. This post covers them too so the three paradigms can be compared in one place.
The CS231N Spring 2026 diffusion lecture doesn't start with DDPM math. It opens by warning that terminology and notation in this area are a mess, then teaches one clean modern version: rectified flow. In training, pick a point between a data sample and noise and have the network predict the velocity from data toward noise; to generate, start from noise and walk backward for about 50 steps. The lecture then stacks on the practical pieces: classifier-free guidance, noise schedules that emphasize middle noise levels, diffusion on VAE latents, Transformers (DiT) as the denoiser, and distillation to cut the step count. Only at the end does it fold VP, VE, and ε/v-prediction into a generalized diffusion framework and name three mathematical views: latent variable model, score function, and SDE.
The CS231N Spring 2026 vision-and-language lecture replaces the "one model per task" approach of the first half of the course with foundation models: pre-train one model on a large, diverse dataset, then adapt it to many tasks through fine-tuning, zero-shot, or few-shot use. Three threads carry the lecture. First, CLIP: contrastive learning in both directions over 400 million image-text pairs scraped from the web, then writing class names as sentences to classify without any fine-tuning; it also has weak spots, such as failing to tell "a mug in some grass" from "some grass in a mug". Second, vision-language models from LLaVA and Flamingo to Qwen3-VL and Molmo, which feed image features into an LLM so it can look at an image and output text. Third, chaining: letting an LLM write descriptions or programs that string existing vision models together.
Assignment 3 in CS231N Spring 2026 is worth 15% of the grade and turns L8, L12–L14 and L16 into four Colab notebooks. Q1 has you write multi-head attention and a Transformer decoder for COCO captioning, then assemble a ViT and train it on CIFAR-10. Q2 implements SimCLR's augmentations and contrastive loss and compares linear classification with and without self-supervised pretraining. Q3 builds DDPM's noising, UNet, denoising loss, sampling and classifier-free guidance to generate text-conditioned 32×32 emoji. Q4 uses pretrained CLIP for similarity, zero-shot classification and retrieval, then segments a video with DINO features trained on a single labeled frame. This guide covers structure, files and targets only. No solutions.
CS231N Lecture 15 runs on one question: what data structure should a 3D shape use so a neural network can read it and produce it? The slides walk through five representations (depth map/surface normals, voxels, point clouds, triangle meshes, implicit surfaces), each with a signature architecture (fully convolutional depth prediction, 3D convolution, PointNet, Pixel2Mesh and Mesh R-CNN, DeepSDF). Then comes the speed trade-off between NeRF and 3D Gaussian Splatting, and a closing roll call of 2025–2026 models: VGGT, TRELLIS, Marble. The 2025 recording uses a different slide deck, with a different order and emphasis.
The last two lectures of CS231N Spring 2026 have no public slides. The schedule lists L17 only as "World Modeling" with guest lecturer Gordon Wetzstein, and L18 only as "Human-Centered AI." Outside readers get 2025 substitutes: that year's L17 was a different topic, Robot Learning (Yunzhu Li, slides and video), and L18 is a Fei-Fei Li recording with no slides. This post labels each year separately and never presents 2025 content as 2026. The second half covers the final project: 35% of the grade, two tracks (Applications and Models), pixels required, and deliverables of a one-paragraph proposal, three milestone check-ins, a 6–8 page report and a poster.