🌏 中文版
Source years: slides and assignments are from Spring 2026; the recordings are from Spring 2025 (YouTube). The two may differ. This post follows the 2026 slides and uses the recording only as a supplement.
This is part 17 of the Reading Stanford CS231N series. The previous post is L14: Generative Models II, Diffusion; the next is A3 guide: Transformer Captioning, SSL, DDPM, CLIP & DINO. L15 (3D vision) moves to after A3 in this series; see the L15 guide.
For the CS231N lecture on May 26, 2026, the schedule gives the title "Vision and Language", while the slide cover reads "Vision + Language (and Foundation Models)". The official material is the 122-page lecture_16.pdf; the matching public recording is Spring 2025 Lecture 16.
The gap between recording and slides is especially large for this lecture. The 2025 lecture_16.pdf has 148 pages, was given by Ranjay Krishna, and includes a long section on Segment Anything; the 2026 version lists Segment Anything only in the taxonomy, and adds Qwen3-VL, SigLIP, and omni models. When the 2025 recording covers something the 2026 slides don't, treat it as extra material.
Scene: how far can one model per task go?
Slide 2 recaps how the course has thought about models so far: train a specialized model for each task. Four data domains, four models, four tasks.
Slide 3 switches to another paradigm: the foundation model. Pre-train one model on a large-scale, diverse dataset, then adapt it to many tasks through fine-tuning, zero-shot, or few-shot use.
How do you tell whether a model counts as a foundation model? Slide 4 gives two tiers:
- Always seen: general and robust across many different tasks
- Often seen: many parameters, lots of data, a self-supervised pre-training objective
Slide 7 sorts foundation models into five classes and marks the ones this lecture covers: Language (ELMo, BERT, GPT, T5), Classification (CLIP, CoCa), LM + Vision (LLaVA, Flamingo, GPT, Gemini, Qwen), Chaining (LMs + CLIP, Visual Programming), and And More (Segment Anything, Whisper, DALL·E, Stable Diffusion, Imagen).
Intuition: widen SimCLR's representation space to hold sentences too
Slides 9–11 recap SimCLR from L12: learn image features with a self-supervised objective that pulls two augmentations of the same image together and pushes different images apart, in the hope that the learned representation generalizes to new instances.
Slide 13 asks the key question: what if this representation space could also embed sentences? If "My favorite dog is a golden retriever" and "A cute fluffy cat" could live in the same space as images, how would you build that joint image-text space?
Mechanism 1: CLIP in two steps
Step 1: collect a ton of data (slides 14–15). CLIP's training data was scraped at scale from web images and their alt-text, about 400 million image-text pairs. The slides note this data collection has since been replicated (Xu et al., Demystifying CLIP Data, ICLR 2024).
Step 2: pick a loss (slides 17–25). The slides quote the CLIP authors on prior work such as image captioning: instead of predicting the exact words for an image, you only need a model that matches the correct description to the image.
So CLIP uses the same family of contrastive objective as SimCLR (InfoNCE). A batch of N images and N texts goes through an image encoder and a text encoder, and pairwise similarities form an N×N matrix. A two-way InfoNCE loss (image-to-text and text-to-image) pushes up the similarities on the diagonal. The slides note that some details aren't shown, such as the temperature parameter and L2 normalization of the vectors. After training you have a model that embeds images and text and gives a similarity score between an image and a text.
Mechanism 2: classify without fine-tuning
The traditional route (slides 26–27) transfers the pre-trained encoder to downstream tasks such as classification, detection, and segmentation with a linear classifier on top.
Language models allow another use (slide 28): no fine-tuning, just use the model "in a creative way", for example by turning review classification into a fill-in-the-blank: "The movie review 'I hated the movie' is ____". Can a vision-language model do the same?
CLIP's clever trick (slides 30–37):
- Use the text encoder to produce a vector for each class, such as "plane", "dog", "bird"
- Run a new image through the image encoder and score it against each class vector
- Pick the most similar class
The slides compare this to 1-NN with the text vectors as the training data. Two prompt-engineering tricks:
| Technique | Gain on ImageNet (per the slides) |
|---|---|
| Write the class as a sentence, "A photo of a [category]", since CLIP was trained on sentences | +1.3% |
| Use several templates ("A photo of…", "A drawing of…") and average the vectors | +5% |
Results (slides 38–43): after training on 400 million image-text pairs, CLIP's zero-shot accuracy matches a ResNet-101 trained on ImageNet — and CLIP used no human labels at all. The generalization story is more interesting: a model trained on ImageNet drops when moved to ObjectNet (same classes, odd viewpoints), while zero-shot CLIP does well, and likewise on graphic images, sketches, and adversarial datasets.
How can no labels beat labels? Slides 46–48 offer three possible answers: "no labels" is a bit misleading (the text is itself supervision), the pre-training scale is massive, and test set leakage. The third probably isn't the main cause: dataset leakage is around 2%. The scale comparison:
| CLIP | ImageNet ResNet | |
|---|---|---|
| Parameters | 307 million | 44.5 million |
| Training data | 400 million images | 1.28 million |
Two common variants (slides 49–52):
- SigLIP: uses a sigmoid instead of a softmax. The benefit is that computing each sample's loss doesn't require materializing the whole batch, which saves memory
- CoCa: adds a generative objective on top of CLIP, in the form of a decoder with a captioning loss
Strengths and weaknesses of CLIP-style models
Strengths (slide 54):
- The dot product is very efficient: easy to train and scale, fast at inference, for example retrieval over 5 billion images
- Open vocabulary, with zero-shot generalization
- Can be chained with other models (CuPL, later in the lecture)
Weaknesses (slides 55–62). The slides cite an April 2022 example from Tristan Thrush et al.: CLIP cannot distinguish "there is a mug in some grass" from "there is some grass in a mug".
- It relies too heavily on batch size to learn concepts. Larger batches give finer concepts: at batch size 4 it learns "animal", at 100 "dog", at 32,000 "Welsh Corgi". But there is a limit: even in a batch of 32K, you are unlikely to see both "a mug in some grass" and "some grass in a mug". This is the compositionality problem, with benchmarks such as Winoground, CREPE, and ARO. One fix is hard-negative fine-tuning ("horse eating grass" vs. "grass eating horse"), but that has its own problem: "a black cat and a brown dog" and "a brown dog and a black cat" describe the same thing — a "hard positive" that should not be pushed apart
- Image-level captions are insufficient supervision. You can also train on region captions with bounding-box coordinates
- No single 5-billion-image dataset contains everything. Data collection and filtering have to be very intentional
Mechanism 3: letting an LLM see — LLaVA and Flamingo
Motivation (slide 65): language models that do next-token prediction can handle a wide range of tasks at inference time — math, sentiment analysis, symbolic reasoning. Can we build a model that takes images and text as input and outputs text? That is a vision-language model (VLM).
Historical context (slides 66–67): VLMs didn't start with LLaVA; they go back at least to ViLBERT in 2019. But those models had to be fine-tuned for each task separately, with non-trivial task-specific methods such as Mask R-CNN bounding-box re-ranking for RefCOCO — the same task-specific paradigm the lecture criticized at the start.
LLaVA's key idea (slides 68–75): an LLM already decodes text autoregressively, so insert image tokens ahead of the text tokens. Which image tokens work best? A CLIP encoder is a good option, but the layer matters:
- The patch tokens in the final layer are not supervised (CLIP's contrastive loss only looks at the CLS or pooling token), so they could be random and the loss wouldn't change
- So use the penultimate layer. In practice these tokens preserve spatial and linguistic information best for the LLM, and dropping CLS gives slight gains
LLaVA's training recipe has three steps: initialize the decoder with a pre-trained LLM (such as LLaMA) and use a pre-trained CLIP as the image encoder; train a new linear layer to bridge CLIP features into the LLM's input space; then fine-tune the LLM and the linear layer together. The slides note that about 178,000 samples of image + instruction + output text give reasonable performance.
Flamingo's alternative fusion (slides 76–85): images go through a vision encoder and feed into the language model's layers from the side, leaving only an <image> tag in the text sequence. The architecture diagram on slide 78 marks the vision encoder and the LM blocks as frozen; the two learned parts are the Perceiver sampler, which converts a variable number of image tokens into a fixed number, and the gated cross-attention layers inserted between LM blocks (labeled GATED XATTN-DENSE in the figure). The training data is arranged like language modeling, with special tags such as <image> and <eos> marking when an image appears or the text ends. It supports in-context learning and works zero-shot and few-shot.
The state of play in 2026: open weights are not fully open source
Who is state of the art? (slide 86) The slides say Gemini is widely considered the best proprietary VLM.
Open models (slides 87–94): the slides first separate open weights (download and run locally) from fully open source (reproduce the training). They then use the Qwen3-VL technical report (December 2025) to show what changed since LLaVA:
- Native image resolution: larger images get more tokens, with 2D-RoPE positional embeddings to handle varying sizes
- A SigLIP-2 vision encoder instead of CLIP
- For video, frame times are given in text, such as
<0.0 seconds> - Embeddings from layers 8, 16, and 24 of SigLIP-2 are used
- Four training phases, starting with bridging the vision encoder and later including context extension
Most open-weight models are distilled (slides 96–102, Molmo and PixMo, CVPR 2025). The slides sort models into API only, open weights, distilled, and completely open, then compare data sizes: Molmo's PixMo has about 700k image-text pairs, against 6 billion for the Llama 3.1V shown on the slide. That is a quality-versus-quantity tradeoff: internet data is often incidental ("pink, japan, aesthetic"), while human annotations are intentional. The catch is that dense captions are hard to collect. Molmo's clever fix: people don't like to type, but they love to talk. Annotators spoke about an image for 60 to 90 seconds, and the speech was automatically converted to text for pre-training.
Mechanism 4: chaining — stringing models together
CuPL (slides 104–107, Pratt et al.): how does a model classify a concept it has never seen (marimba, viaduct, papillon, lorikeet)? First have an LLM generate a description of the class, then classify using the description.
VisProg (slides 108–116, Gupta & Kembhavi): can chaining generalize to all vision tasks? The traditional options are to train a new model for the new task, or hand-write a Python script that chains existing models (say, run single-image VQA twice and combine the answers) — but the hand-written script only covers two images. VisProg has GPT generate that program, composing off-the-shelf vision models to answer the question.
Omni models (slide 117): beyond vision and language, train one model to take in and output text, audio, and video. The slides say this started with GPT-4o in 2024, and that Thinking Machines and Gemini have recently introduced their own.
Back to the models: where this lecture sits in the course
Looking back, this lecture gathers parts from earlier ones: the ViT from L8 is CLIP's image encoder, the contrastive learning of L12 becomes image-text contrast, and the captioning of L7 is the path CLIP's authors decided not to take. The FLUX.1 model from the previous lecture also uses CLIP as one of its text encoders.
Assignment mapping. Q4 of A3 is CLIP and DINO (CLIP_DINO.ipynb, cs231n/clip_dino.py). The CLIP part reuses the COCO data from the A2 captioning task and asks you to implement image-text similarity, a zero-shot classifier, and a text-to-image retriever; one inline question is weakness #1 from this lecture: "Why does CLIP's learning depend on the batch size?" Details are in the A3 guide.
Going deeper
Official material
- lecture_16.pdf (2026)
- lecture_16.pdf (2025): includes the Segment Anything section that matches the 2025 recording
- Recording: Spring 2025 L16
- A3: Q4 CLIP & DINO
Related posts on this site (each is self-contained; overlap is kept)
- CS224N: Multimodality — the same family of models from an NLP course
- CS336 Lecture 17: Multimodal alignment — image tokens from the language-model training side
Access limits
Under the grading in the global AI/CS course map, this course is A3: the 2026 slides, assignments, and starter code are public, plus full 2025 recordings. The gaps for this lecture: 2026 recordings are on Canvas for enrolled students only, and the 2025 recording differs from the 2026 slides more than in other lectures.
References
- Stanford CS231N course homepage (Spring 2026)
- CS231N Spring 2026 schedule
- CS231N Spring 2026 Lecture 16 slides
- CS231N Spring 2025 schedule
- CS231N Spring 2025 Lecture 16 slides
- YouTube: CS231N Spring 2025 Lecture 16: Vision and Language
- CS231N Assignment 3 (Spring 2026)
- Radford et al., Learning Transferable Visual Models From Natural Language Supervision (CLIP)
- Zhai et al., Sigmoid Loss for Language Image Pre-Training (SigLIP)
- Yu et al., CoCa: Contrastive Captioners are Image-Text Foundation Models
- Liu et al., Visual Instruction Tuning (LLaVA)
- Alayrac et al., Flamingo
- Qwen3-VL Technical Report
- Deitke et al., Molmo and PixMo
- Pratt et al., What does a platypus look like? (CuPL)
- Gupta & Kembhavi, Visual Programming (VisProg)
Loading...