Skip to content

CMU 10-423 L14–L15: Cross-Attention, DiT, Prompt-to-Prompt, and Q-Former

Sep 30, 20261 min
TL;DRWhere does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.

🌏 中文版

This post is based on the Spring 2026 offering of CMU 10-423/623/723 Generative AI. It's post 14 in the Reading CMU 10-423 series. The main materials are the Lecture 14 slides (plus an inked version) and the Querying Transformer first half of the Lecture 15 slides. The scaling-laws second half is in post 16.

The recordings live on CMU's Panopto and aren't available outside CMU, so this post relies on the slides alone. The Q-Former part of L15 is almost entirely paper figures with no slide text, so the description below comes from the captions of the BLIP-2, MetaQueries, and Perceiver IO figures the slides reproduce. The schedule lists no readings for these two lectures. I checked every fact against the official materials on 2026-09-30.

The previous post said LDM reads the prompt through cross-attention without taking it apart. This post answers: where in the model does the text condition enter, and why can editing the prompt change only part of the image?

Cross-attention: queries and keys/values come from different places

L14 starts by reviewing scaled dot-product attention from L2: one input sequence x is multiplied by W_q, W_k, and W_v to get queries, keys, and values. Cross-attention changes one thing. Queries come from one sequence y (length n), while keys and values come from another sequence x (length m).

Matrix form
  • Q = Y·W_q ∈ ℝ^{n×d}
  • K = X·W_k ∈ ℝ^{m×d}
  • V = X·W_v ∈ ℝ^{m×d}
  • S = QKᵀ/√d ∈ ℝ^{n×m}
  • A = softmax(S), Y′ = AV

The slides start with translation. m is the number of source-language tokens ("estoy llegando tarde") and n is the number of target-language tokens ("I am running late"). Each target word's attention weights form a probability distribution over the source words.

Now switch to LDM:

  • Queries come from a layer of the UNet
  • Keys and values come from the text encoder's representation of the prompt
  • m is the number of prompt tokens; n is the number of latent dimensions, or the number of pixels if there's no compression
  • Each (latent) pixel's attention weights form a probability distribution over the prompt tokens

The slides add one note: the real attention and cross-attention blocks are multi-head.

This view, in which every pixel has a distribution over the words it looks at, is the foundation for Prompt-to-Prompt later on.

Classifier-free guidance: making generation follow the prompt

The slides frame a trade-off. Diffusion models, unlike GANs, are good at producing diverse samples, but once you add a condition, that diversity can pull generations away from the prompt. Classifier-free guidance (CFG) pulls them back.

The same noise model learns two modes that share parameters: a conditional ε_θ(z_t, t, c), and ε_θ(z_t, t, ∅) with the condition replaced by a null embedding ∅. Sampling combines the two:

The sampling algorithm on the slide
  1. w = 7.5
  2. c = tokenize("a cat with green eyes"), c∅ = tokenize("")
  3. z_T ∼ N(0, I)
  4. At each step: ε_θ ← (1 + w)·ε_θ(z_t, t, c) − w·ε_θ(z_t, t, c∅), then compute z_{t−1} with the DDPM update

The rule on the slide: a larger w makes samples match the condition more closely, at the cost of diversity. The next two pages show figures from Ho & Salimans in which samples look more and more like their class as the guidance scale rises from 0 to 3.

Diffusion Transformer: the UNet is optional

The slides' timeline starts with the UNet for medical image segmentation in 2015 and runs through DDPM and Dhariwal & Nichol, all using UNets. Everyone kept using it "because it seems to work well". DiT showed you don't actually need the UNet.

The DiT backbone is essentially a ViT with a few tweaks (see L5):

  • Input: a noisy latent, a timestep, and a class label (or other conditioning)
  • Output: a mean and a covariance of fixed size
  • After a final LayerNorm, a linear layer turns the T token embeddings into the fixed-size output

How the condition goes in: adaLN-Zero

The slides call the conditioning mechanism the interesting part of the DiT block. The original paper tried four options:

  1. In-context conditioning
  2. A cross-attention block
  3. Adaptive LayerNorm (adaLN)
  4. adaLN with zero initialization (adaLN-Zero)

adaLN-Zero worked best empirically. The key idea is to learn an MLP that outputs the scale and shift parameters for LayerNorm and the residual connections.

Note the contrast with LDM. LDM reads the condition through cross-attention; DiT's best variant uses adaLN. That difference matters in HW4.

Measuring scale in GFLOPs, not parameters

The slides define two metrics. GFLOPs measure the compute of one forward pass, independent of hardware. FID passes real and generated images through a pretrained Inception-v3, takes mid-layer features, fits a multivariate Gaussian to each set, and computes the Fréchet distance between the two Gaussians. Lower is better.

Shrinking DiT's patch size raises GFLOPs without adding parameters, so DiT studies FID against GFLOPs. The slides leave a question here: if the patch size drops from 4 to 2, how much does total compute grow?

The scaling results come in three points:

  1. GFLOPs and FID are strongly correlated: more compute, better FID
  2. For the same training compute, larger DiT models are more compute-efficient than smaller ones
  3. At similar compute, DiT also beats an LDM with a UNet backbone (the slides compare LDM-4 and DiT-XL/2)

Prompt-to-Prompt: change the prompt, change only that part

Why it's needed

The slides list the problems with two older approaches:

  • Fix the random seed and change the prompt: the simplest baseline, but the whole layout can change dramatically. It feels like generating an unrelated image, not editing
  • Mask-based editing (e.g. Blended Diffusion): a mask marks which region stays fixed and the text decides how the rest changes, but the user has to draw the mask

Prompt-to-Prompt aims to edit with text alone, no mask.

How it works

It assumes a pretrained LDM and involves no training at all; it only changes how samples are drawn:

  1. Encode the original prompt y, run diffusion once, and record the attention weights A_{T−1}, …, A_1 at every step
  2. Encode the edited prompt y* and run diffusion again:
    • Reuse the initial noise z_T from the first run
    • Until timestep τ, use the attention weights from the first run
    • After that, switch to the attention weights computed in this run
    • Whichever weights you use, you still attend to y*
  3. If running in latent space, decode back to pixel space at the end

The slides pose a question here: why use the original attention weights for a while before switching to the new ones? The answer is written on the inked version, and I didn't transcribe it. You can work it out against the Prompt-to-Prompt results slide that varies the switching point.

When the shapes don't match

If y and y* have different lengths, A_t and A*_t won't have the same shape. The slides' fix is to swap in only the matching parts:

  • The latent dimension stays constant
  • With a fixed-length text encoder (e.g. CLIP's 77 positions, padded with <PAD>), the prompt dimension is also constant
  • But the words may not line up. The slides' example changes "orange cat sitting" to "big tabby cat sitting" and copies the attention weights for "orange" to both "big" and "tabby"

The results slide explains that word-level cross-attention swapping automatically finds which regions of the image should stay fixed and which should change. Beyond word swaps, Prompt-to-Prompt supports down-weighting a descriptor and inserting phrases to change style or content. Each type of edit maps to a different cross-attention manipulation.

First half of L15: the Querying Transformer

Starting from PaliGemma's linear projection

L15 opens by recalling PaliGemma from the previous post: SigLIP image embeddings pass through one linear projection into the LLM, and the LLM has to learn how to use them on its own. The next three slides show a different way to connect the two.

BLIP-2's Q-Former

BLIP-2 places a lightweight Querying Transformer between a frozen image encoder and a frozen LLM, and pretrains it in two stages.

Stage 1: vision-language representation learning. The Q-Former takes a set of learnable queries. The queries run self-attention among themselves and read the image encoder's output through cross-attention (in every other block). Three objectives are optimized jointly, each with a different self-attention mask that controls how queries and text interact:

ObjectiveMask
Image-Text MatchingBidirectional
Image-Grounded Text GenerationMultimodal causal
Image-Text Contrastive LearningUnimodal

Stage 2: vision-to-language generative learning. The Q-Former's output queries pass through a fully connected layer that maps them to the LLM's input dimension, and they go into the frozen LLM ahead of the text. A decoder-only LLM (e.g. OPT) generates text directly. With an encoder-decoder LLM (e.g. FlanT5), the prefix text goes to the encoder and the decoder generates the suffix.

The slides follow with a full page of zero-shot instructed image-to-text examples.

MetaQueries and Perceiver IO

  • MetaQueries: the slide title is "Simpler yet effective". Learnable queries attach directly to a frozen multimodal LLM, pull out conditions for generation, and pass them through a connector to a diffusion model. Training uses only a denoising objective on paired data. The figure's tagline reads "Render unto diffusion what is generative, and unto LLMs what is understanding"
  • Perceiver IO: labeled a "Historical Note". A smaller latent array reads an input of any size through cross-attention, does most of its computation in latent space, and decodes to an output of any size with an output query array

What the three share: a fixed number of learnable queries that pull the needed information out of a variable-length input from another modality.

This is what HW4 asks you to build

HW4's programming part (40 points) ties L14 and L15 together. The setup in hw4.pdf:

  • GPT-2 (frozen): turns text into per-token features
  • DiT (frozen): a class-conditional diffusion transformer pretrained on CIFAR-10, which originally reads a class embedding through adaLN
  • Q-Former (trained): 2–8 learnable queries read GPT-2's hidden states through cross-attention and output per-layer conditioning vectors that replace DiT's original class embedding

The starter code already implements CFG, and training drops the text condition with probability 0.1. The assignment ends by having you inspect cross-attention maps to see what the Q-Former pulls out of the LLM. Details are in the HW4 post.

Assessment

The schedule puts Quiz 4 on March 16, covering the text-to-image part of L12 through L15. The questions aren't available.

How to read these two lectures

  1. Match the five-line matrix form of cross-attention to the translation example before reading the LDM slide. Once you're clear on what n and m stand for, Prompt-to-Prompt's shape problem is intuitive.
  2. When reading DiT, put adaLN-Zero side by side with LDM's cross-attention. Both inject a condition: one adjusts LayerNorm's parameters, the other lets the image read the text.
  3. The Q-Former slides are figures only, so open Figure 2 of the BLIP-2 paper and go through the three masks cell by cell.

One thing to do tonight: write down what Prompt-to-Prompt's Edit(A_t, A*_t, t) returns when t ≥ τ and when t < τ. Then think about the slide's question: what would happen to the layout if you used the new attention weights from the very first step?

Further reading

Series navigation: previous L12–L13: Text-to-image, latent diffusion, and vision-language models | next HW4: Text-to-image with a Q-Former | Series overview

References