Skip to content

CS231N Assignment 3: Transformer Captioning, Self-Supervised Learning, DDPM, CLIP and DINO

Sep 30, 20261 min
TL;DRAssignment 3 in CS231N Spring 2026 is worth 15% of the grade and turns L8, L12–L14 and L16 into four Colab notebooks. Q1 has you write multi-head attention and a Transformer decoder for COCO captioning, then assemble a ViT and train it on CIFAR-10. Q2 implements SimCLR's augmentations and contrastive loss and compares linear classification with and without self-supervised pretraining. Q3 builds DDPM's noising, UNet, denoising loss, sampling and classifier-free guidance to generate text-conditioned 32×32 emoji. Q4 uses pretrained CLIP for similarity, zero-shot classification and retrieval, then segments a video with DINO features trained on a single labeled frame. This guide covers structure, files and targets only. No solutions.

🌏 中文版

Which year: The assignment follows the CS231N Spring 2026 Assignment 3 page and the assignment3.zip starter code (downloaded 2026-09-30; the notebooks were last modified in May 2026). Recordings for the related lectures come from the Spring 2025 YouTube playlist. The 2026 recordings are on Canvas for enrolled students only, and the two years may differ.

This is post 18 in the Reading Stanford CS231N series.

Series: previous L16: Vision and Language | next L15: 3D Vision | series overview

A3 is the last of CS231N's three assignments. Attention, self-supervision, diffusion and CLIP have lived on slides so far. This assignment makes you write them as code that runs and passes numeric checks.

The four questions span the course's third unit:

QuestionNotebookMain lectures
Q1 Image Captioning with TransformersTransformer_Captioning.ipynbL8 Attention and Transformers
Q2 Self-Supervised Learning for Image ClassificationSelf_Supervised_Learning.ipynbL12 Self-Supervised Learning
Q3 Denoising Diffusion Probabilistic ModelsDDPM.ipynbL14 Generative Models 2, with background in L13
Q4 CLIP and DINOCLIP_DINO.ipynbL16 Vision and Language (CLIP), L12 (DINO)

This post covers only what the official page and the starter code show: questions, files and targets. It gives no solutions.

The basics

The assignments page makes A3 15% of the course grade. The 2026 schedule releases it on May 14 (the day of L13) and sets the deadline at 11:59pm Pacific on May 28.

The assignment page lists four goals: understand and implement Transformers and combine them with CNN features for captioning; use self-supervised learning to help image classification; implement DDPM for image generation; implement and understand CLIP and DINO. Most of the code is PyTorch.

The official workflow is Google Colab. Local development isn't officially supported, but a requirements.txt is included if you want your own virtual env. When the four notebooks are done, you run collect_submission.ipynb. It zips your code into a3_code_submission.zip and converts all notebooks into one a3_inline_submission.pdf, and both go to Gradescope.

Every notebook opens with a "Student Declaration" asking whether and how you used generative AI. The assignments page treats generative AI like a collaborator: you write your solutions independently, and using it to substantially complete the work violates the Honor Code. The same page says the staff knows past solutions are online and still expects your own work.

Q1: from captioning to ViT, two Transformers in one notebook

The notebook picks up from A2's RNN captioning. You already captioned COCO images with a vanilla RNN; now you do the same job with a Transformer decoder. Unlike A2, this one is mostly PyTorch, not NumPy.

Everything you fill in lives in cs231n/transformer_layers.py and cs231n/classifiers/transformer.py, in this order:

  1. MultiHeadAttention: multi-head scaled dot-product attention, with dropout on the attention weights. The notebook spells out that each head projects to d/h dimensions and the scale factor is 1/√(d/h).
  2. PositionalEncoding: sine/cosine position codes added to the word embeddings.
  3. TransformerDecoderLayer: self-attention, cross-attention over image features, and a per-position feedforward block.
  4. CaptioningTransformer.forward: assemble the captioning model, overfit it on a small dataset (the notebook wants a final training loss below 0.05), then compare against the RNN using the provided sampling code.

The second half switches to the Vision Transformer. You fill in PatchEmbedding, TransformerEncoderLayer and the VisionTransformer forward pass, then tune architecture and hyperparameters to exceed 0.45 test accuracy on CIFAR-10 after 2 epochs. This ViT average-pools all patch vectors before classifying and reuses the same 1D sinusoidal positional encoding.

Every step comes with a numeric check, with tolerances from e-3 to 1e-6. You move on only when it passes.

There are three inline questions. Why multiple heads, why divide by √(d/h), and why add a linear layer after attention? Why do ViTs often lose to CNNs on small datasets, and what helps? And how does self-attention cost change if you separately double the hidden dimension, the image side length, the patch size, or the number of layers?

Think about that last one before you code. It checks whether you've absorbed that token count depends on image size and patch size.

Q2: SimCLR, and whether pretraining actually helps

The assignment page reminds you to switch the Colab runtime to GPU before you start.

The notebook uses SimCLR as its template. Augment the same image twice, pass both through an encoder f to get representations h, then through a projection head g to get z. The contrastive loss pulls the two z's of the same image together. After training, g is discarded and only f is kept for downstream tasks.

The work sits in cs231n/simclr/:

  • data_utils.py: compute_train_transform() applies a random resized crop to 32×32, a horizontal flip with probability 0.5, color jitter with probability 0.8, and grayscale with probability 0.2. CIFAR10Pair.__getitem__() returns the augmented pair.
  • contrastive_loss.py: first sim and a pairwise simclr_loss_naive, then the vectorized sim_positive_pairs, compute_sim_matrix and simclr_loss_vectorized. All checks use a 1e-7 tolerance. The notebook notes that it reorders positive pairs in the batch, so its indices differ from the paper's; the change makes vectorization easier.
  • utils.py: complete train() using your vectorized loss.

You don't train from scratch. The course provides weights pretrained on CIFAR-10 for about 18 hours, and the notebook has you load them and train a bit more (about 10 minutes).

Then comes the experiment the whole question builds to. Remove the projection head, freeze every earlier layer, and train only a linear classifier. Compare its test accuracy with a baseline that had no self-supervised pretraining and trains all weights, and plot the two. The notebook warns up front that the baseline will look low but reasonable.

Q2 has no inline question. What you hand in is that comparison plot.

Q3: DDPM, from the noising equation to text-conditioned generation

The notebook trains a DDPM to generate 32×32 emoji conditioned on text prompts. Text goes through a pretrained CLIP text encoder into a 512-dimensional vector. To save time, the training-set text is already encoded.

The steps follow the equations in the DDPM paper:

  1. q_sample (cs231n/gaussian_diffusion.py): the forward noising process from Eq. (4). The check expects zero relative error.
  2. predict_start_from_noise and predict_noise_from_start: the model can predict either the clean image or the noise, and each can be recovered from the other.
  3. Unet.__init__ and Unet.forward (cs231n/unet.py): define the down- and upsampling blocks and the forward pass, taking x_t, t and the text embedding and returning a tensor of the same shape.
  4. p_losses: the denoising training objective.
  5. p_sample: one reverse-process sampling step, from Eq. (6). The outer sample loop is already written.
  6. Unet.cfg_forward: classifier-free guidance. The notebook gives the update ε ← (w+1)·ε(x_t, t, c) − w·ε(x_t, t, ∅), with the condition dropped at some rate during training.

The training code is in cs231n/ddpm_trainer.py and needs no changes. The rest of the notebook samples from a pretrained model in cs231n/exp/pretrained. You can train your own on a Colab GPU, but the notebook says it may take more than 12 hours on a T4 before generations look reasonable.

It is also candid about limits. The emoji dataset is small, not enough for a text-to-image model that generalizes. Unseen prompts may come out poorly, and a higher guidance scale doesn't guarantee faithfulness to the text.

Q3 also has no inline question.

Q4: CLIP and DINO, two representations learned without labels

The CLIP half extracts features for COCO images and captions with a pretrained CLIP (the captioning data again, but matching pairs instead of generating). Then, in cs231n/clip_dino.py, you implement:

  • get_similarity_no_loop: the text–image similarity matrix, without loops.
  • clip_zero_shot_classifier: zero-shot classification from one description per class. The notebook lists the 10 expected predictions so you can compare.
  • CLIPImageRetriever: the reverse direction, retrieving images from text. Comments give the expected top two for each query.

The DINO half starts with the motivation. Contrastive methods like SimCLR and CLIP need very large batches. BYOL avoids many negatives with a student–teacher setup, and DINO follows that idea: the student learns by backpropagation, while the teacher skips backprop and tracks an exponential moving average of the student's weights.

Three steps follow. Visualize the attention maps from the [CLS] token to each patch in the last layer of a DINO ViT. Run PCA on patch features and color the image by the top three components. Finally, implement DINOSegmentation: train a lightweight per-patch classifier on the labels of a single frame from one DAVIS video, then segment the rest of that video. The targets are mean IoU above 0.45 on the first test frame, above 0.50 on the last, and above 0.55 over the whole video. The notebook suggests a linear layer or a two-layer MLP with suitable weight decay to avoid overfitting.

There are five inline questions. Why does CLIP's learning depend on batch size, and what can you do if batch size is fixed? How would you extend image–text alignment to more than two modalities? Where do the tensor shapes printed in the attention-map section come from? What structure does the PCA view show, and what do same-colored and differently colored regions mean? And would a segmentation model trained on CLIP ViT patch features beat DINO or lose to it, and why?

That last question ties L12 to L16. Neither CLIP nor DINO uses human labels, but their training objectives teach patch features different things.

What's in the starter code

assignment3.zip is about 24 MB, most of it one GIF used in the DINO demo. The files that matter for the four questions:

PathUsed in
cs231n/transformer_layers.py, cs231n/classifiers/transformer.pyQ1 layers and models to fill in
cs231n/captioning_solver_transformer.py, cs231n/classification_solver_vit.pyQ1 training loops (provided)
cs231n/simclr/ (data_utils.py, contrastive_loss.py, utils.py, model.py)Q2
cs231n/gaussian_diffusion.py, cs231n/unet.py, cs231n/ddpm_trainer.py, cs231n/emoji_dataset.pyQ3
cs231n/clip_dino.pyQ4 functions and classes to fill in
cs231n/datasets/get_datasets.sh and friendsDownload COCO and other data

The zip still contains images/styles/ and a style-transfer example image, but none of the four 2026 questions uses style transfer.

What a self-learner can get

Available: the assignment page, the full questions and numeric checks in all four notebooks, the .py skeletons, and the download steps for the Q2 and Q3 pretrained weights. (The SimCLR weight URL in the Q2 notebook still returned HTTP 200 on 2026-09-30; Q3 uses trainer.download_pretrained(), which I haven't run.) By design, a Colab GPU is enough to get every question to "checks pass, results visible." Under the scheme in the global AI course map, that is access level A3.

Not available: the Gradescope autograder and the rubric for the inline questions, the TA explanations on Ed, and the 2026 lecture recordings. The numeric checks only confirm that your implementation matches the reference. Nobody grades your inline answers.

How to self-study it

  1. Go Q1 → Q2 → Q3 → Q4. Q1's attention and ViT are what let you read DINO's attention maps in Q4. Q2's contrastive loss is what explains why CLIP cares about batch size.
  2. Read the matching lecture slides before each question. When a check fails, go back to the equations rather than searching for past solutions.
  3. Write your inline answers in two or three sentences, then test them. For Q1's third question, actually double the image side length and time the attention.
  4. Q2's comparison plot and Q4's IoU numbers are the most research-like parts. Treat them as a warm-up for the final project.

One thing to do tonight: open Transformer_Captioning.ipynb, read only the four scenarios in Inline Question 3, and write down on paper how token count and compute change in each.

Further reading

References