🌏 中文版
Source years: The slides are the Spring 2026 Lecture 12 slides from CS231N (104 pages, cover date 2026-05-07). The recording is the Spring 2025 Lecture 12 on YouTube (about 1 hour 14 minutes; the 2025 schedule lists Ehsan Adeli as lecturer). The 2026 recordings are on Canvas for enrolled students only, so the two years may differ.
This is part 14 of the Reading Stanford CS231N series and the first lecture of the course's third unit, "Generative and Interactive Visual Intelligence."
The lecture starts from something the course already showed. In the 4096-dimensional features from AlexNet's last layer, nearest neighbors are semantically similar images, while nearest neighbors in pixel space are not. Learned representations are useful, but they need a lot of labeled data. Can we train such representations without huge manually labeled datasets?
The framework: pretext and downstream tasks
The slides split self-supervised learning into two stages:
- Pretext task: a task defined from the data itself, with no manual annotation (arguably a form of unsupervised learning). Labels are generated automatically. Training yields an encoder
- Downstream task: attach the encoder to the task you actually care about and train it with a small amount of labeled data (supervised or semi-supervised)
We usually don't care how well the pretext task goes. We care how useful the learned features are for downstream classification, detection, and segmentation.
How to evaluate. The slides list several angles: performance on the pretext task itself; representation quality, including the linear evaluation protocol (freeze the encoder, train a linear classifier), clustering, and t-SNE visualization; robustness and generalization to other datasets; computational efficiency; and transfer to downstream tasks.
The slides also zoom out: the same idea shows up in language modeling (GPT-4), speech synthesis (WaveNet), and robot learning.
Part 1: labels from image transformations
| Pretext task | What the model does | Source (as cited on the slides) |
|---|---|---|
| Rotation prediction | Rotate the whole image by 0/90/180/270 degrees; 4-way classification | Gidaris et al. 2018 |
| Relative patch location | Given two patches, predict where the second sits relative to the first | Doersch et al. 2015 |
| Jigsaw puzzles | Put shuffled patches back in order | Noroozi & Favaro 2016 |
| Inpainting | Remove a region and reconstruct the missing pixels with an encoder-decoder | Pathak et al. 2016 (Context Encoders) |
| Colorization | Predict color from grayscale; the split-brain autoencoder uses two sub-networks | Zhang et al. |
| Video colorization | Color the other frames to match a colored reference frame | Vondrick et al. 2018 |
A few details worth keeping:
- The rotation hypothesis: a model can recognize the correct rotation only if it has the "visual commonsense" of what the object should look like unperturbed. The evaluation freezes the first two conv layers on CIFAR-10 and trains the later layers on a subset of labels.
- The inpainting loss is reconstruction plus an adversarial term: the slides compare results side by side for reconstruction only, adversarial only, and both (adversarial learning is the topic of the next lecture).
- Tracking emerges from video colorization: the task relies on colors staying consistent over time. The slides show that the learned attention can propagate segmentation masks across frames, even though nobody taught the model to track.
MAE: push the mask ratio to 75%
Masked Autoencoders (He et al., 2021) are the representative reconstruction method:
- Split the image into non-overlapping patches as in ViT and uniformly mask a very large fraction (75%)
- The encoder sees only the unmasked 25%: linear projection, positional embeddings, Transformer blocks. Because its input is small, the encoder can be very large, with over 9× the computation per token of the decoder
- The decoder merges the encoder output with a shared mask token in the masked positions, adds positional encodings, runs Transformer blocks, and projects back to pixels
- The loss is MSE in pixel space, computed only on masked patches
A high mask ratio makes the task hard enough to be meaningful. The slides list MAE's long list of ablations: mask ratio, decoder depth and width, whether the encoder uses mask tokens, reconstruction target, data augmentation, mask sampling, and training schedule.
The limitation of this part: the learned representation may be tied to one specific pretext task. Is there a more general one?
Part 2: contrastive learning
The more general idea: different versions of the same object should have nearby representations, and different objects should be pushed apart (attract and repel).
Formally, x is a reference sample, x⁺ a positive, and x⁻ a negative. Choose a score function and learn an encoder f that scores positive pairs (x, x⁺) high and negative pairs (x, x⁻) low.
The InfoNCE loss (standard form)
With one positive and N−1 negatives:
$$\mathcal{L} = -\mathbb{E}\left[\log \frac{\exp(s(f(x), f(x^+)))}{\exp(s(f(x), f(x^+))) + \sum_{j=1}^{N-1} \exp(s(f(x), f(x_j^-)))}\right]$$
This is an N-way cross-entropy: pick the positive out of N candidates. The slides' summary notes that it is a lower bound on the mutual information between f(x) and f(x⁺).
SimCLR
SimCLR (Chen et al., 2020):
- Cosine similarity as the score function
- Positives come from data augmentation: augment the same image twice at random (random cropping, color distortion, blur) to get a positive pair; the other images in the batch are negatives
- A projection network g(·) sits on top of the features, and the contrastive loss is computed in the projected space
The slides add a note: the assignment uses a slightly different SimCLR formulation, so follow the assignment instructions.
Two key design choices:
- A non-linear projection head helps. The slides' possible explanation: the contrastive objective makes the representation invariant to augmentations, which may discard information useful downstream. The projection head lets the z space carry that invariance, so the h space before it keeps more information
- Large batches are crucial. But they have a large memory footprint during backprop, and the ImageNet experiments needed distributed training on TPUs (the subject of the previous lecture)
For evaluation, the slides show a linear classifier trained on frozen SimCLR features from ImageNet, and semi-supervised results from fine-tuning the encoder with only 1% or 10% of ImageNet labels.
MoCo and MoCo v2
MoCo (He et al., 2020) removes SimCLR's dependence on huge batches. The key differences:
- Keep a FIFO queue of keys as negatives
- Compute gradients and update the encoder only through the queries; the key encoder gets no gradient and is updated by momentum
- As a result, minibatch size is decoupled from the number of negatives, so you can use many negatives
MoCo v2 is a hybrid: the non-linear projection head and strong augmentation from SimCLR, plus MoCo's momentum-updated queue, which allows many negatives without TPUs. The slides' takeaway is that a non-linear projection head and strong data augmentation are crucial for contrastive learning.
CPC: contrast at the sequence level
SimCLR and MoCo are instance-level methods (positives and negatives are different instances). CPC (Contrastive Predictive Coding, van den Oord et al., 2018) works at the sequence level:
- Contrastive: contrast the "right" sequence with "wrong" ones
- Predictive: predict future patterns from the current context
- Coding: first encode each sample in the sequence into a vector z_t
The slides show an audio example (linear classification on LibriSpeech) and an image example: split an image into patches, treat rows of patches as a top-to-bottom sequence, and use the upper rows to predict the lower ones. The summary slide's verdict: CPC applies to many problems, but it is less effective for image representations than instance-level methods.
Part 3: DINO, self-distillation with no labels
The slides end with DINO (Caron et al., 2021, Emerging Properties in Self-Supervised Vision Transformers), which is also a suggested reading on the schedule. The slides are mostly figures, so the mechanism below follows the original paper:
- Two networks, one architecture: a student g_s and a teacher g_t. The student is updated with SGD. The teacher receives no gradient (stop-gradient); its parameters are an exponential moving average of the student's, which makes it a momentum encoder
- Objective: pass both outputs through a softmax to get probability distributions, and use cross-entropy to make the student match the teacher. The paper interprets this as knowledge distillation with no labels
- Multi-crop: each image yields 2 global views at 224² and several local views at 96². The student sees every view; the teacher sees only the global ones, which encourages local-to-global correspondence
- Avoiding collapse: DINO uses no negatives, only centering and sharpening of the teacher's output. The paper explains that centering stops one dimension from dominating but pushes toward a uniform output, while sharpening does the opposite; balancing the two, together with a momentum teacher, is enough to avoid collapse
The DINO paper's pseudocode skeleton (without multi-crop)
# gs, gt: student and teacher networks
# C: center (K)
# tps, tpt: student and teacher temperatures
# l, m: network and center momentum rates
gt.params = gs.params
for x in loader: # load a minibatch x with n samples
x1, x2 = augment(x), augment(x) # random views
s1, s2 = gs(x1), gs(x2) # student output n-by-K
t1, t2 = gt(x1), gt(x2) # teacher output n-by-K
loss = H(t1, s2)/2 + H(t2, s1)/2
loss.backward() # back-propagate
# student, teacher and center updates
update(gs) # SGD
gt.params = l*gt.params + (1-l)*gs.params
C = m*C + (1-m)*cat([t1, t2]).mean(dim=0)
def H(t, s):
t = t.detach() # stop gradient
s = softmax(s / tps, dim=1)
t = softmax((t - C) / tpt, dim=1) # center + sharpen
return - (t * log(s)).sum(dim=1).mean()
The paper's abstract lists two main findings: self-supervised ViT features explicitly contain the semantic segmentation of an image, which is much less clear in supervised ViTs and convnets; and these features are also excellent k-NN classifiers. It also stresses the importance of the momentum encoder, multi-crop training, and small patches. The last DINO slide is titled "DINO v2," but it is figures only, so this post does not describe it.
How to study it
- Get the pretext/downstream split straight before the methods. For every new method, ask two things: where do the labels come from automatically, and how is the representation evaluated?
- InfoNCE is just cross-entropy. Think of it as a pick-1-of-N classification, and SimCLR, MoCo, and CPC differ only in where positives and negatives come from.
- Do A3. In the Spring 2026 assignment covered in the A3 guide, Q2 is Self-Supervised Learning (the starter code has a
simclr/folder) and Q4 is CLIP & DINO. Remember the slides' warning that the assignment's SimCLR formulation differs slightly from the slides. - Gaps: the schedule lists this lecture's topics as pretext tasks, contrastive learning, and multisensory supervision, but neither the 2026 nor the 2025 slide agenda has a separate multisensory section; the CPC audio example is the closest. For audio-visual self-supervision, go back to the audio-visual part of L10. The 2026 recording is not public.
Further reading
- Contrastive learning extended to image-text pairs (CLIP): L16: Vision and Language in this series
- ViT and patch basics: L8: Attention, Transformers, and ViT in this series
- Suggested reading from the schedule: Lilian Weng's Self-Supervised Representation Learning
Series navigation: Previous: L11: Large-Scale Distributed Training | Next: L13: Generative Models I: VAEs, GANs, and Autoregressive Models | Series overview
References
- CS231N course homepage (Spring 2026)
- CS231N Spring 2026 schedule
- Lecture 12: Self-Supervised Learning slides (Spring 2026, PDF)
- Lecture 12 slides (Spring 2025, PDF)
- CS231N Spring 2025 schedule
- Stanford CS231N 2025 Lecture 12: Self-Supervised Learning (YouTube)
- Stanford CS231N 2025 YouTube playlist
- CS231N Assignment 3 (Spring 2026)
- Caron et al. (2021). Emerging Properties in Self-Supervised Vision Transformers (DINO)
- Meta AI blog: DINO and PAWS
- Lilian Weng (2019). Self-Supervised Representation Learning
- Gidaris et al. (2018). Unsupervised Representation Learning by Predicting Image Rotations
- Pathak et al. (2016). Context Encoders: Feature Learning by Inpainting
- He et al. (2021). Masked Autoencoders Are Scalable Vision Learners
- Chen et al. (2020). A Simple Framework for Contrastive Learning of Visual Representations (SimCLR)
- He et al. (2020). Momentum Contrast for Unsupervised Visual Representation Learning (MoCo)
- van den Oord et al. (2018). Representation Learning with Contrastive Predictive Coding
Loading...