Skip to content

CS231N L1: Where Computer Vision Came From, and Where This Course Is Going

Sep 30, 20261 min
TL;DRThe first CS231N lecture of 2026 comes in two slide decks. The first tells the history of vision and deep learning on a single timeline: Hubel & Wiesel's cat experiments, Marr's stages of visual representation, then the Neocognitron, backprop, and LeNet, until ImageNet and AlexNet join the two threads. The second covers the course map, grading, and rules, and moves every assignment onto Colab. After this lecture you'll know which gap each of the remaining 17 lectures fills.

🌏 中文版

Source years: slides are from Spring 2026; for a recording, see Spring 2025 Lecture 1 on YouTube. They may differ; this post follows the 2026 slides. This is post 1 of the Reading Stanford CS231N series. For the course's positioning, grading, access gaps, and the 10-week plan, see the series overview.

The first CS231N lecture of 2026 was on March 31, taught by Fei-Fei Li and Ehsan Adeli. The schedule links two decks for it: part 1 is a short history of computer vision and deep learning, and part 2 is the course overview and rules.

There are no equations in this lecture. Its job is to show why the course starts with image classification, and why it goes all the way to generative models, 3D, and world modeling.

One Venn diagram to place the course

Early in part 1, a Venn diagram grows one circle at a time (the slides credit Justin Johnson for the idea). Artificial Intelligence comes first, then Machine Learning inside it, then Computer Vision and Deep Learning. Finally the slide labels "This class": the overlap of computer vision and deep learning.

Around the edges sit mathematics, neuroscience, physics, psychology, biology, and computer science. The point is that the course borrows ideas from many fields but follows one main line: deep learning applied to vision problems.

An earlier slide shows three words, Big Data, Neural networks, and GPUs, grouped as the "Modern AI Revolution". Those three words foreshadow the whole history that follows.

Thread one: how computer vision became a recognition problem

The slides start far back, with the Cambrian explosion 530–540 million years ago ("Evolution's Big Bang"), then the camera obscura. Computer vision proper comes next. In the order of the slide timeline:

  • 1959, Hubel & Wiesel: recordings from cat brains showed simple cells that respond to edges at particular orientations and complex cells that respond to orientation and movement, with some translation invariance
  • 1963, Larry Roberts: "Machine Perception of Three-Dimensional Solids", which went from the original picture to an edge image to selected feature points
  • 1970s, David Marr: visual representation in stages, from primal sketch to 2½-D sketch to 3-D model
  • 1970s: recognition via parts, such as generalized cylinders and pictorial structures
  • 1986, Canny, and 1987, Lowe: recognition via edge detection
  • 1997, Normalized Cuts (Shi & Malik): recognition via grouping
  • 1999, SIFT (David Lowe): recognition via matching
  • 2001, Viola & Jones face detection: the slides call it one of the first successful applications of machine learning to vision
  • PASCAL VOC and Caltech101: standardized recognition datasets and challenges appear

An "AI winter" runs through the middle: expert systems failed to deliver, funding dried up, but subfields such as vision, NLP, and robotics kept growing. Meanwhile, cognitive science and neuroscience work continued. The slides cite rapid serial visual perception (RSVP) experiments and a 1996 Nature paper by Thorpe et al. marked "150 ms !!", and land on one conclusion: "Visual recognition is a fundamental task for visual intelligence."

That sentence explains where the course starts. After decades of vision research, the focus settled on recognition, so L2 begins with image classification.

Thread two: how neural networks fell out of favor and came back

On the same timeline, the slides layer in the history of neural networks:

  • 1958, Perceptron (Rosenblatt)
  • 1969, Minsky & Papert: showed that perceptrons can't learn XOR, which, per the slides, caused a lot of disillusionment
  • 1980, Neocognitron (Fukushima): directly inspired by Hubel & Wiesel, it interleaves simple cells (convolution) and complex cells (pooling), but had no practical training algorithm
  • 1986, Backprop (Rumelhart, Hinton, Williams): successfully trained multi-layer perceptrons
  • 1998, LeNet (LeCun et al.): applied backprop to a Neocognitron-like architecture to recognize handwritten digits, and NEC deployed it to process checks. The slides call it "very similar to our modern convolutional networks"
  • "Deep Learning" from 2006: people tried to train deeper and deeper networks, but the slides add one line: "No good dataset to work on"

That last line is where the threads meet. The algorithms existed. The data didn't.

Where they meet: ImageNet and AlexNet

The slides then list dataset sizes: Caltech101 at about 9K images, LabelMe at 37K, PASCAL VOC at 30K, SUN at 131K, and then ImageNet at 15 million images in 22,000 categories (Deng et al., CVPR 2009). The ImageNet classification challenge uses 1,000 of those classes and 1,431,167 images.

The next node on the timeline is AlexNet in 2012. From there, the slides use two charts, ImageNet top-1 accuracy climbing year after year and submissions to the top computer vision conference exploding, to make the case for "2012 to Present: Deep Learning Explosion". Part 2 adds that the original 2009 ImageNet paper won the IEEE PAMI Longuet-Higgins Prize at CVPR 2019, an award for the most influential computer vision paper from ten years earlier.

A dozen or so "Deep Learning is Everywhere" slides follow. They roughly preview the rest of the course:

What the slides showMatching lecture
AlexNet (2012) → ResNet (2015) → ViT (2021) → DiT (2023)L5, L6, L8, L14
Faster R-CNN detection, Segment Anything segmentationL9
Video understanding and activity recognitionL10
Early image captioning (2015) next to a detailed 2026 caption from Gemini 3L7, L16
GANs, DALL·E, and later image generationL13, L14

The last slides turn to limits and responsibility. Computer vision "still has a long way to go". It can cause harm, through stereotypes and hiring decisions, and it can save lives, through ambient sensing in hospitals and home care. The slides split visual intelligence into Understanding, Reasoning, and Generation, and close on Spatial Intelligence and World Modeling.

The course map: the official materials don't cut it the same way

The overview slide in part 2 splits the course into four blocks:

  1. Deep Learning Basics
  2. Perceiving and Understanding the Visual World
  3. Generative and Interactive Visual Intelligence
  4. Human-Centered Applications and Implications

The closing slide of the same deck says something else: Deep Learning Basics (L2–4), Perceiving and Understanding the Visual World (L5–12), Reconstructing and Interacting with the Visual World (L13–17), and Human-Centered Artificial Intelligence (L18). The official schedule has only three units: L2–L4, L5–L11, and L12–L18.

The three versions differ on where L12 (self-supervised learning) belongs and whether L18 stands alone. This series follows the schedule's three units and covers L18 in the final post.

The examples part 2 gives for each block map onto the later lectures:

  • Basics: image classification → linear classifiers → regularization and optimization → neural networks
  • Perceiving and understanding: tasks beyond classification (semantic segmentation, object detection, instance segmentation, video, visualization), models beyond the MLP (CNNs, RNNs, Transformers), and large-scale distributed training (data parallelism, model parallelism, synchronous vs. asynchronous updates)
  • Generative and interactive: self-supervised learning, style transfer and image generation, vision-language models (with CLIP's contrastive pre-training as the example), 3D vision, and embodied intelligence. The slides say A3 has you implement a generative model that produces emojis from text

Course rules: Colab, the Honor Code, and a new midterm

The second half of part 2 is logistics. What matters for self-learners:

  • Every assignment runs on Google Colab. A1 goes out 4/2 and is due 4/16. It covers kNN, a Softmax linear classifier, a two-layer network, image features, and deeper networks with optimizers
  • Every assignment has a coding part and a written part. Code is autograded; the written part is graded by TAs
  • Grading: three assignments at 12% + 18% + 15% = 45%, a midterm at 20% (marked "New" on the slide), the project at 35%, and up to 3% extra credit for participation
  • Four Honor Code rules: don't look at solutions or code that aren't yours (including from AI tools), don't share your solution code, credit anyone you worked with, and "Do not submit AI-generated responses"
  • Learning objectives: formalize vision applications as tasks; learn to code, debug, and train CNNs; use frameworks such as PyTorch and TensorFlow; understand where the field stands and what ethical questions come before deployment

The optional reading lists three free books: Deep Learning by Goodfellow, Bengio, and Courville; Mathematics of Deep Learning (the slides point to chapters 5–7 for vector calculus and continuous optimization); and Dive into Deep Learning.

What to do tonight

If you plan to follow the series, do two things tonight:

  1. Open the Python/Numpy tutorial on cs231n.github.io and make sure broadcasting and vectorized operations make sense. The kNN question in A1 asks you to compute distances without loops
  2. Redraw the part 1 timeline yourself, and next to each node write the later lecture it connects to. L2 immediately uses the "data-driven" idea, which is exactly what the ImageNet part of the story concludes

The 2026 recordings are on Canvas only and closed to outside readers. If you want to listen along, Spring 2025 Lecture 1 is on YouTube and runs about an hour. I did not compare it slide by slide with the 2026 deck.

Further reading

Series navigation: Previous: Overview and self-study plan | Next: L2: Image classification, kNN, and linear classifiers

References