Skip to content
All tags

#clip

7 posts

CMU 10-423 L12–L13: Text-to-Image, Latent Diffusion, and Vision-Language Models

CMU 10-423 spends two lectures connecting generative models to a second modality. The second half of L12 asks how text can steer an image: three routes (GANs, autoregressive Parti, diffusion with DALL-E 2 and Imagen) lead to latent diffusion, which compresses images into an autoencoder's latent space, runs DDPM there, and reads the prompt through cross-attention. L13 goes the other way and lets a language model read images: CLIP/SigLIP or a VQ-VAE turns the image into vectors or integers for a decoder-only Transformer. What separates read-only VLMs (PaliGemma, Qwen-VL) from VLMs that can also output images (LWM, Gemini) is whether image tokens are discrete.

CS231N Assignment 3: Transformer Captioning, Self-Supervised Learning, DDPM, CLIP and DINO

Assignment 3 in CS231N Spring 2026 is worth 15% of the grade and turns L8, L12–L14 and L16 into four Colab notebooks. Q1 has you write multi-head attention and a Transformer decoder for COCO captioning, then assemble a ViT and train it on CIFAR-10. Q2 implements SimCLR's augmentations and contrastive loss and compares linear classification with and without self-supervised pretraining. Q3 builds DDPM's noising, UNet, denoising loss, sampling and classifier-free guidance to generate text-conditioned 32×32 emoji. Q4 uses pretrained CLIP for similarity, zero-shot classification and retrieval, then segments a video with DINO features trained on a single labeled frame. This guide covers structure, files and targets only. No solutions.

CS231N L16: Vision and Language — From CLIP's Contrastive Learning to Multimodal Foundation Models That Talk About Images

The CS231N Spring 2026 vision-and-language lecture replaces the "one model per task" approach of the first half of the course with foundation models: pre-train one model on a large, diverse dataset, then adapt it to many tasks through fine-tuning, zero-shot, or few-shot use. Three threads carry the lecture. First, CLIP: contrastive learning in both directions over 400 million image-text pairs scraped from the web, then writing class names as sentences to classify without any fine-tuning; it also has weak spots, such as failing to tell "a mug in some grass" from "some grass in a mug". Second, vision-language models from LLaVA and Flamingo to Qwen3-VL and Molmo, which feed image features into an LLM so it can look at an image and output text. Third, chaining: letting an LLM write descriptions or programs that string existing vision models together.

Reading NCCU Yen-Lung Tsai Generative AI, L11: Text-to-Image AI, Principles and Practice — CLIP Reads the Prompt, Schedulers Decide Whether It Converges, LoRA Learns Only ΔW, and You Build a Web App with diffusers

L11 fills in the rest of the Stable Diffusion diagram. CLIP is trained so that matching text and images get similar vectors, which turns a prompt into a 77×768 embedding. Schedulers compress 1,000 noising steps into twenty or thirty denoising steps, but ancestral samplers such as Euler a never settle: push to 100 steps and the subject changes jackets and seats. LoRA freezes the original W and learns only a ΔW factored into A·B. The hands-on part loads an SD 1.5-family model with diffusers, and the week 11 homework is your own image-generation web app.

Reading NTU ADL 2025 Fall: Beyond Supervised Learning and Multimodality — Auto-Encoders, VAE, Dual Learning, Contrastive Learning, and CLIP

Big data is not big annotated data. The last ADL lecture asks how to learn good representations without labels, and answers: find the latent factors that control the data. An auto-encoder squeezes the input into a short code and reconstructs it. The denoising version adds noise or masks 15% of tokens first, which is exactly the idea behind BERT's masked LM. A VAE forces the code to follow a distribution, so you can sample from it to generate. Dual learning lets paired tasks, such as translation and back-translation or understanding and generation, act as feedback for each other. Self-supervised learning has two camps: self-prediction (hide part, guess it back) and contrastive learning (pull similar pairs together, push dissimilar ones apart). CLIP runs contrastive learning on 400 million image-text pairs, making zero-shot image classification possible, and DALL·E 2 uses CLIP's representations to generate images. Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.

CS336 Lecture 17: Multimodal Models Turn Images into Tokens, Then Reconcile Semantics with Detail

Lecture 17 organizes CLIP/SigLIP, LLaVA, Qwen-VL, and Chameleon into three paths: contrastive encoders learn semantics, vision-encoder/projector/LM stacks provide understanding, and discrete image tokens enable generation. Resolution, token budgets, and modality balance constrain them all.

Multimodal RAG: Bringing Images into the Knowledge Base

Climbing routes carry a ton of visual information (topos, wall photos) that text-only RAG misses entirely. Multimodal RAG makes images searchable and understandable.