Skip to content

CS231N L12: Self-Supervised Learning — Learning Good Representations Without Labels

Sep 30, 20261 min
TL;DRCS231N Lecture 12 asks whether we can learn good representations without huge manually labeled datasets. The answer comes in three parts. First, pretext tasks that generate labels from image transformations: predicting rotation, solving jigsaw puzzles, inpainting, colorization, and MAE with a 75% mask ratio. Second, the more general contrastive learning: the InfoNCE loss, SimCLR with its large batches, MoCo, which decouples batch size from the number of negatives with a queue, and sequence-level CPC. Third, DINO, which needs no negatives: a student predicts the output of a momentum teacher, and centering plus sharpening prevent collapse. The core evaluation is linear probing: freeze the encoder and train only a linear classifier.

🌏 中文版

Source years: The slides are the Spring 2026 Lecture 12 slides from CS231N (104 pages, cover date 2026-05-07). The recording is the Spring 2025 Lecture 12 on YouTube (about 1 hour 14 minutes; the 2025 schedule lists Ehsan Adeli as lecturer). The 2026 recordings are on Canvas for enrolled students only, so the two years may differ.

This is part 14 of the Reading Stanford CS231N series and the first lecture of the course's third unit, "Generative and Interactive Visual Intelligence."

The lecture starts from something the course already showed. In the 4096-dimensional features from AlexNet's last layer, nearest neighbors are semantically similar images, while nearest neighbors in pixel space are not. Learned representations are useful, but they need a lot of labeled data. Can we train such representations without huge manually labeled datasets?

The framework: pretext and downstream tasks

The slides split self-supervised learning into two stages:

  1. Pretext task: a task defined from the data itself, with no manual annotation (arguably a form of unsupervised learning). Labels are generated automatically. Training yields an encoder
  2. Downstream task: attach the encoder to the task you actually care about and train it with a small amount of labeled data (supervised or semi-supervised)

We usually don't care how well the pretext task goes. We care how useful the learned features are for downstream classification, detection, and segmentation.

How to evaluate. The slides list several angles: performance on the pretext task itself; representation quality, including the linear evaluation protocol (freeze the encoder, train a linear classifier), clustering, and t-SNE visualization; robustness and generalization to other datasets; computational efficiency; and transfer to downstream tasks.

The slides also zoom out: the same idea shows up in language modeling (GPT-4), speech synthesis (WaveNet), and robot learning.

Part 1: labels from image transformations

Pretext taskWhat the model doesSource (as cited on the slides)
Rotation predictionRotate the whole image by 0/90/180/270 degrees; 4-way classificationGidaris et al. 2018
Relative patch locationGiven two patches, predict where the second sits relative to the firstDoersch et al. 2015
Jigsaw puzzlesPut shuffled patches back in orderNoroozi & Favaro 2016
InpaintingRemove a region and reconstruct the missing pixels with an encoder-decoderPathak et al. 2016 (Context Encoders)
ColorizationPredict color from grayscale; the split-brain autoencoder uses two sub-networksZhang et al.
Video colorizationColor the other frames to match a colored reference frameVondrick et al. 2018

A few details worth keeping:

  • The rotation hypothesis: a model can recognize the correct rotation only if it has the "visual commonsense" of what the object should look like unperturbed. The evaluation freezes the first two conv layers on CIFAR-10 and trains the later layers on a subset of labels.
  • The inpainting loss is reconstruction plus an adversarial term: the slides compare results side by side for reconstruction only, adversarial only, and both (adversarial learning is the topic of the next lecture).
  • Tracking emerges from video colorization: the task relies on colors staying consistent over time. The slides show that the learned attention can propagate segmentation masks across frames, even though nobody taught the model to track.

MAE: push the mask ratio to 75%

Masked Autoencoders (He et al., 2021) are the representative reconstruction method:

  • Split the image into non-overlapping patches as in ViT and uniformly mask a very large fraction (75%)
  • The encoder sees only the unmasked 25%: linear projection, positional embeddings, Transformer blocks. Because its input is small, the encoder can be very large, with over 9× the computation per token of the decoder
  • The decoder merges the encoder output with a shared mask token in the masked positions, adds positional encodings, runs Transformer blocks, and projects back to pixels
  • The loss is MSE in pixel space, computed only on masked patches

A high mask ratio makes the task hard enough to be meaningful. The slides list MAE's long list of ablations: mask ratio, decoder depth and width, whether the encoder uses mask tokens, reconstruction target, data augmentation, mask sampling, and training schedule.

The limitation of this part: the learned representation may be tied to one specific pretext task. Is there a more general one?

Part 2: contrastive learning

The more general idea: different versions of the same object should have nearby representations, and different objects should be pushed apart (attract and repel).

Formally, x is a reference sample, x⁺ a positive, and x⁻ a negative. Choose a score function and learn an encoder f that scores positive pairs (x, x⁺) high and negative pairs (x, x⁻) low.

The InfoNCE loss (standard form)

With one positive and N−1 negatives:

$$\mathcal{L} = -\mathbb{E}\left[\log \frac{\exp(s(f(x), f(x^+)))}{\exp(s(f(x), f(x^+))) + \sum_{j=1}^{N-1} \exp(s(f(x), f(x_j^-)))}\right]$$

This is an N-way cross-entropy: pick the positive out of N candidates. The slides' summary notes that it is a lower bound on the mutual information between f(x) and f(x⁺).

SimCLR

SimCLR (Chen et al., 2020):

  • Cosine similarity as the score function
  • Positives come from data augmentation: augment the same image twice at random (random cropping, color distortion, blur) to get a positive pair; the other images in the batch are negatives
  • A projection network g(·) sits on top of the features, and the contrastive loss is computed in the projected space

The slides add a note: the assignment uses a slightly different SimCLR formulation, so follow the assignment instructions.

Two key design choices:

  • A non-linear projection head helps. The slides' possible explanation: the contrastive objective makes the representation invariant to augmentations, which may discard information useful downstream. The projection head lets the z space carry that invariance, so the h space before it keeps more information
  • Large batches are crucial. But they have a large memory footprint during backprop, and the ImageNet experiments needed distributed training on TPUs (the subject of the previous lecture)

For evaluation, the slides show a linear classifier trained on frozen SimCLR features from ImageNet, and semi-supervised results from fine-tuning the encoder with only 1% or 10% of ImageNet labels.

MoCo and MoCo v2

MoCo (He et al., 2020) removes SimCLR's dependence on huge batches. The key differences:

  • Keep a FIFO queue of keys as negatives
  • Compute gradients and update the encoder only through the queries; the key encoder gets no gradient and is updated by momentum
  • As a result, minibatch size is decoupled from the number of negatives, so you can use many negatives

MoCo v2 is a hybrid: the non-linear projection head and strong augmentation from SimCLR, plus MoCo's momentum-updated queue, which allows many negatives without TPUs. The slides' takeaway is that a non-linear projection head and strong data augmentation are crucial for contrastive learning.

CPC: contrast at the sequence level

SimCLR and MoCo are instance-level methods (positives and negatives are different instances). CPC (Contrastive Predictive Coding, van den Oord et al., 2018) works at the sequence level:

  • Contrastive: contrast the "right" sequence with "wrong" ones
  • Predictive: predict future patterns from the current context
  • Coding: first encode each sample in the sequence into a vector z_t

The slides show an audio example (linear classification on LibriSpeech) and an image example: split an image into patches, treat rows of patches as a top-to-bottom sequence, and use the upper rows to predict the lower ones. The summary slide's verdict: CPC applies to many problems, but it is less effective for image representations than instance-level methods.

Part 3: DINO, self-distillation with no labels

The slides end with DINO (Caron et al., 2021, Emerging Properties in Self-Supervised Vision Transformers), which is also a suggested reading on the schedule. The slides are mostly figures, so the mechanism below follows the original paper:

  • Two networks, one architecture: a student g_s and a teacher g_t. The student is updated with SGD. The teacher receives no gradient (stop-gradient); its parameters are an exponential moving average of the student's, which makes it a momentum encoder
  • Objective: pass both outputs through a softmax to get probability distributions, and use cross-entropy to make the student match the teacher. The paper interprets this as knowledge distillation with no labels
  • Multi-crop: each image yields 2 global views at 224² and several local views at 96². The student sees every view; the teacher sees only the global ones, which encourages local-to-global correspondence
  • Avoiding collapse: DINO uses no negatives, only centering and sharpening of the teacher's output. The paper explains that centering stops one dimension from dominating but pushes toward a uniform output, while sharpening does the opposite; balancing the two, together with a momentum teacher, is enough to avoid collapse
The DINO paper's pseudocode skeleton (without multi-crop)
# gs, gt: student and teacher networks
# C: center (K)
# tps, tpt: student and teacher temperatures
# l, m: network and center momentum rates
gt.params = gs.params
for x in loader:  # load a minibatch x with n samples
    x1, x2 = augment(x), augment(x)  # random views
    s1, s2 = gs(x1), gs(x2)  # student output n-by-K
    t1, t2 = gt(x1), gt(x2)  # teacher output n-by-K
    loss = H(t1, s2)/2 + H(t2, s1)/2
    loss.backward()  # back-propagate
    # student, teacher and center updates
    update(gs)  # SGD
    gt.params = l*gt.params + (1-l)*gs.params
    C = m*C + (1-m)*cat([t1, t2]).mean(dim=0)

def H(t, s):
    t = t.detach()  # stop gradient
    s = softmax(s / tps, dim=1)
    t = softmax((t - C) / tpt, dim=1)  # center + sharpen
    return - (t * log(s)).sum(dim=1).mean()

The paper's abstract lists two main findings: self-supervised ViT features explicitly contain the semantic segmentation of an image, which is much less clear in supervised ViTs and convnets; and these features are also excellent k-NN classifiers. It also stresses the importance of the momentum encoder, multi-crop training, and small patches. The last DINO slide is titled "DINO v2," but it is figures only, so this post does not describe it.

How to study it

  • Get the pretext/downstream split straight before the methods. For every new method, ask two things: where do the labels come from automatically, and how is the representation evaluated?
  • InfoNCE is just cross-entropy. Think of it as a pick-1-of-N classification, and SimCLR, MoCo, and CPC differ only in where positives and negatives come from.
  • Do A3. In the Spring 2026 assignment covered in the A3 guide, Q2 is Self-Supervised Learning (the starter code has a simclr/ folder) and Q4 is CLIP & DINO. Remember the slides' warning that the assignment's SimCLR formulation differs slightly from the slides.
  • Gaps: the schedule lists this lecture's topics as pretext tasks, contrastive learning, and multisensory supervision, but neither the 2026 nor the 2025 slide agenda has a separate multisensory section; the CPC audio example is the closest. For audio-visual self-supervision, go back to the audio-visual part of L10. The 2026 recording is not public.

Further reading

Series navigation: Previous: L11: Large-Scale Distributed Training | Next: L13: Generative Models I: VAEs, GANs, and Autoregressive Models | Series overview

References