Skip to content

CS231N L10: Video Understanding — What Changes When You Add a Time Axis

Sep 30, 20261 min
TL;DRCS231N Lecture 10 treats video as 2D plus time, a T×3×H×W tensor, and follows one thread: efficiency. Train on short clips and average several clips at test time. Architectures run from per-frame 2D CNNs and late fusion to 3D CNNs, then two-stream networks that isolate motion with optical flow, and I3D, which inflates 2D weights into 3D. After 2021 the field moved to Transformers, where token counts explode, which led to divided space-time attention, Video Swin, MViT, and tubelets. The last part covers temporal localization, audio-visual models, VideoLLMs, and long-form video, where HourVideo shows how far the field still has to go.

🌏 中文版

Source years: The slides are the Spring 2026 Lecture 10 slides from CS231N (92 pages, cover date 2026-04-30). The recording is the Spring 2025 Lecture 10 on YouTube (about 1 hour 8 minutes; the 2025 schedule lists Ruohan Gao as lecturer). The 2026 recordings are on Canvas for enrolled students only, so the two years may differ.

This is part 12 of the Reading Stanford CS231N series.

The previous post moved the output from one label per image to every object and every pixel. This lecture extends in a different direction: the input gains a time axis. The slides set it up in one line: a video is a sequence of images, a 4D tensor of shape T×3×H×W.

One question ties the lecture together: where in the network do you handle the extra T, and at what cost? Each architecture is a different answer.

Task and data: from objects to actions

Image classification recognizes dogs, cats, and trucks. Video classification recognizes swimming, running, jumping, and eating. The example dataset is Sports-1M: 1 million YouTube videos labeled with 487 sports (Karpathy et al., CVPR 2014).

The first obstacle is size. Video usually runs at 30 fps. Uncompressed at 3 bytes per pixel, SD (640×480) takes about 1.5 GB per minute and HD (1920×1080) about 10 GB per minute. That does not fit in GPU memory.

The course's fix is practical:

  • Training: classify short clips at low fps
  • Testing: run the model on several clips from the same video and average the predictions

Every architecture that follows assumes short clips.

First-generation answers: how CNNs handle time

ApproachWhere time is fusedWhat the slides say
Single-frame CNNNever; classify each frame independently and average probabilities at test timeOften a very strong baseline
Late fusion (FC)Run a 2D CNN per frame, flatten T×D×H'×W', feed an MLPCaptures high-level appearance per frame; how do you handle arbitrary length?
Late fusion (pooling)Run a 2D CNN per frame, average-pool over space and timeGood for high-level scene information, bad at comparing low-level motion between frames
Early fusion / 3D CNN3D convolution and pooling from the first layer, fusing time graduallyEvery activation is a 4D tensor, D×T×H×W

Try single-frame first. This is the most practical line in the section. Before a video project, run a plain 2D CNN on each frame, average, and treat that as the baseline to beat.

How 3D convolution differs from 2D. The slides go through it piece by piece. The filter gains a time dimension (for example 2×3×5×5). The sliding window works the same way with one more direction. You can stride in time. Each activation map becomes T'×28×28. Nothing is conceptually new; it is the convolution from L5 pushed one dimension further.

Two tricks that improve results

Trick 1: pull motion out on its own (two-stream)

The slides cite Johansson's 1973 biological-motion experiment: people recognize an action from a few moving points of light. Motion alone carries a lot of information.

The tool for measuring motion is optical flow, a displacement field F(x, y) = (dx, dy) between two frames with I_{t+1}(x+dx, y+dy) = I_t(x, y). You can compute it with a classical algorithm or approximate it faster with a neural network.

Two-stream networks (Simonyan & Zisserman, NeurIPS 2014) separate appearance from motion:

  • Spatial stream: a single image, 3×H×W
  • Temporal stream: stacked optical flow, [2×(T−1)]×H×W, with the first 2D conv processing all flow images (early fusion)

In the slides' UCF-101 bar chart, the temporal stream alone (83.7) beats the spatial stream alone (73). Fusing both with an SVM reaches 88, against 65.4 for the 3D CNN baseline.

The slides add a note on the present: optical flow is now rarely used directly for video understanding. It survives mostly as an intermediate representation for other tasks such as robot control (citing Ko et al., ICLR 2024).

Trick 2: inflate a 2D network into 3D (I3D)

Image architectures have years of research behind them. Can video reuse them? I3D (Carreira & Zisserman, CVPR 2017):

  1. Take a 2D CNN and replace each K_h×K_w conv or pool layer with a K_t×K_h×K_w 3D version
  2. Initialize the 3D weights from the 2D ones: copy them K_t times along time and divide by K_t
Why divide by K_t

Feed in a "constant" video, the same image repeated K_t times. The 3D convolution sums K_t identical results along time, and dividing by K_t gives exactly the original 2D output. The inflated network therefore starts out behaving like the ImageNet-pretrained 2D network and learns temporal information from there.

The slides show Kinetics-400 top-1 accuracy, all with an Inception backbone, in two groups (trained from scratch and ImageNet-pretrained). The comparison covers a per-frame CNN, CNN+LSTM, two-stream, an inflated 3D CNN, and a two-stream inflated 3D CNN.

Second-generation answers: Transformers and the token explosion

The slides split the field into two eras: 2014–2021 for 3D CNNs plus RNNs, 2021–2026 for Transformers.

The problem with switching to ViT is arithmetic. A 224×224 image cut into 16×16 patches gives 14×14 = 196 tokens. Video blows that up:

InputTokensThe slides' comparison
One image196
16-frame clip3,136About a book chapter
5 minutes at 1 fps58.8kAbout a short novel
5 minutes at 24 fps1,411,200Near current models' context limits

Self-attention cost grows with the square of the token count. The slides give two broad strategies: change the attention operator, or reduce the number of tokens.

Strategy 1: change attention

  • Divided space-time attention (TimeSformer, Bertasius et al., ICML 2021) splits joint space-time attention into two steps. In the time step, each token attends only to the same spatial location in other frames. In the space step, it attends only to other tokens in the same frame. Per-token cost drops from O(NT) to O(N+T), and information still spreads across space and time over many blocks.
  • Video Swin Transformer (CVPR 2022) restricts self-attention to small local space-time cubes. It looks a lot like a 3D CNN with attention inside each cube. The cubes shift between layers so information crosses cube boundaries.
  • MViT (Multiscale Vision Transformers, ICCV 2021) uses convolution to aggregate and shorten the K and V sequences before attention, while the output length stays the same. Like a ResNet, the network halves the spatial size and doubles the channels as it goes, for example from a 56×56 grid at 96 dimensions to 14×14 at 384.

Strategy 2: fewer tokens (tubelets)

ViViT (Arnab et al., ICCV 2021) replaces patches with tubelets, 3D blocks spanning several frames. The slides give two intuitions: patches contain no motion information and tubelets do, and tubelets produce far fewer tokens (spanning 4 frames cuts the count by 4×). ViViT, VideoMAE, Video Swin, MViT, and V-JEPA all use them. Other token-reduction methods (adaptive token selection, token merging, learned compression) are deferred to L16.

Beyond classification: localization, multimodality, long video

So far everything classifies short clips. The second half extends in three directions.

Temporal and spatio-temporal localization. Temporal action localization finds the frames for each action in a long, untrimmed video. It can reuse a Faster R-CNN-style design: generate temporal proposals, then classify them. Spatio-temporal detection goes further: detect every person in space and time and classify what they are doing. The example dataset is AVA.

Audio-visual models. The slides use the McGurk effect (the same sound paired with different lip movements is heard as "Ba" or "Fa") to show that vision changes what we hear. Then several lines of research:

  • Visually guided speech separation: VisualVoice (Gao et al., CVPR 2021) separates a speech mixture into the left and right speakers
  • Musical instrument separation: Gao & Grauman (ICCV 2019) train on 100,000 unlabeled multi-source clips and then separate audio in new videos
  • Audio-visual fusion for action recognition, and audio as a cheap preview to speed up recognition in long videos
  • Multimodal egocentric video, such as Ego-Exo conversational graph prediction (Jia et al., CVPR 2024)

VideoLLMs and long video. The slides list Video-LLaVA, VideoLLaMA 3, and Video-ChatGPT, then ask whether current vision systems can understand long-form video. The test is to ask questions about the video, such as "Where did I leave my AirPods after working out?" in a 1 hour 10 minute egocentric recording. The examples come from HourVideo, with question types like "identify the unique individuals the camera wearer interacted with" and "how can the camera wearer get to the backyard from the kitchen," posed as multiple-choice questions. The slides' takeaway: there is a significant gap in long-form video understanding, and a lot of work left to do.

How to study it

  • Write tensor shapes on paper. Every architecture here differs in which dimension holds T and at which layer it gets fused. For each diagram, write the input, intermediate activation, and output shapes. That beats memorizing model names.
  • Learn to compute token counts yourself. The product 196 × T explains why every video Transformer design is about saving tokens or saving attention.
  • Gaps: this lecture has no dedicated assignment, and none of A1–A3 includes a video problem. The 2026 recording is not public, so for the spoken explanation you need the 2025 recording.

Further reading

Series navigation: Previous: L9: Object Detection, Segmentation, and Visualization | Next: L11: Large-Scale Distributed Training | Series overview

References