Skip to content

CMU 10-423 HW4: Text-to-Image with a Q-Former Between a Frozen GPT-2 and a Frozen DiT — Structure, Files to Edit, and Compute

Sep 30, 20261 min
TL;DRHW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.

🌏 中文版

This guide is based on the Spring 2026 edition of CMU 10-423/623/723 Generative AI. It is part 15 of Reading CMU 10-423 and closes the multimodal unit. The previous post, L14–L15: Cross-Attention, DiT, Prompt-to-Prompt, and the Q-Former, showed where text conditioning enters the model. This one looks at how the homework makes you turn a diffusion model that only knows 10 classes into one that reads text.

Official materials used: hw4.zip from the Coursework page (the 30-page "S26 10423 HW4.pdf", starter code, and unit tests), the read-only Overleaf template, the course schedule, and the Querying Transformer section of the Lecture 15 slides. I downloaded and checked all of them on 2026-09-30.

This post covers structure, points, files to edit, compute, and environment. It gives no solutions.

The basics

ItemDetails
NameHomework 4: Multi-Modal Foundation Models
ScopeL12–L14 (the schedule says "HW4 out (L12-L14)")
ReleasedThe handout says 2026-03-13; the schedule marks "HW4 out" on March 12
Due2026-03-23 (Slot A); the schedule lists March 30 for Slot B (tentative)
SubmissionGradescope: one written PDF; code uploads are train_qformer.py, dit.py, image_caption_data.py
Total79

Points table (handout, page 1):

SectionPoints
LaTeX Template Alignment0
Latent Diffusion Model (LDM)7
VQ-VAEs8
CLIP4
VLMs18
Programming: Text-to-Image Generation40
Code Upload0
Collaboration Questions2

The series overview explains the Slot A (human work only) and Slot B (AI allowed) rules. The reminder slide in L18 marks HW4 Slot A as "no AI assistance!" and announces a Programming Test on HW3/HW4 on March 27.

What the written questions ask

The four written sections map onto the four topics of L12–L14. Most items are short calculations or short answers:

  • LDM (7 points): compute the latent size of a 512×512 image at compression factor f = 8; explain why diffusion runs in latent space rather than pixel space; then change cross-attention so the values come from the image side, and explain what breaks mathematically and why it makes no sense intuitively.
  • VQ-VAE (8 points): take the objective with and without stop-gradient, differentiate each with respect to the encoder output and the codebook vector, and describe in words what stop-gradient does to the two terms and what the codebook is being optimized to do.
  • CLIP (4 points): show that the CLIP paper's symmetric cross-entropy objective is equivalent to raising the cosine similarity of matched image-caption pairs while lowering it for every other pair.
  • VLM (18 points): using PaliGemma2, draw the data flow between the SigLIP vision encoder, the linear projection, and the Gemma 2 language model and describe each part's role; work out how the number of visual tokens changes when resolution goes from 224×224 to 448×448 with 14×14 patches; and finally, ask a proprietary VLM (the handout names GPT or Gemini) for a dense caption, then find an adversarial image-question pair it gets completely wrong, with a screenshot.

That last item is the only one that needs an outside service. You can still do it today, but the result depends on the model version you hit.

Programming: train only the middle layer

Section 6 of the handout sets up a pipeline with three parts:

ComponentStateRole
GPT-2FrozenTurns text into per-token 768-dimensional hidden states
DiT (Diffusion Transformer)FrozenA class-conditional diffusion model pretrained on CIFAR-10; its 10 class embeddings modulate every layer through adaLN
Q-FormerTrainedUses a few learnable queries to cross-attend over GPT-2's output and produce the conditioning vector each DiT layer needs, replacing the class embedding

The handout gives two reasons not to feed GPT-2 features straight into the DiT. GPT-2's output has variable length while the DiT takes a single 768-dimensional vector, and there is no reason for the two representation spaces to line up. The Q-Former design is credited to Perceiver IO, with BLIP-2 and MetaQueries as models that use the same structure. The Querying Transformer section of L15 follows the same order: a PaliGemma recap, BLIP-2's two stages, MetaQueries, then Perceiver IO.

The outputs are 32×32 images. The handout warns that CIFAR-10 is low resolution to begin with, so blurry samples do not mean the model is broken.

Three TODO functions

The training loop, cross-attention, inference, and visualization are all provided. You fill in three functions:

FileFunctionWhat it does
image_caption_data.pyprompts_to_padded_hidden_statesFor each prompt, call the tokenizer and GPT-2 (output_hidden_states=True), keep the requested layer's hidden states, pad to a common length, and return a boolean mask marking non-padding positions
train_qformer.pysetup_optimizer_and_schedulerFreeze every parameter outside the Q-Former; build AdamW with betas fixed at (0.9, 0.95); add a linear warmup through LambdaLR when warmup_steps > 0; use MSE loss
dit.pyQueryEmbedder.forwardLook up the query table, expand it to the batch, run pre-norm self-attention, cross-attention, and FFN residuals block by block, then project with output_proj and query_to_layer to (batch, DiT layers, conditioning dim)

The first function hides an easy miss. GPT-2 has 12 layers, but there is a hidden state before the first layer and after the last, so 13 layer indices are valid. The unit tests also require an error when the index is below 0 or out of range.

The training command in run_in_cloud.ipynb passes --gpt2_layer_index 12, the state after the last layer. The Q-Former defaults to 2 blocks and 8 heads, with 4 queries. The training script builds the DiT with dim 256, 10 layers, and 8 heads.

Fourteen unit tests

test_all.py holds 14 tests, each weighted 1. One checks the submitted files, six cover hidden-state extraction (layer selection, tokenizer before GPT-2, padding, two kinds of invalid index, masks), four cover the optimizer setup, and three cover the Q-Former forward pass (output shape, with and without an attention mask). The tests use a fake GPT-2 and tokenizer, so no weights need downloading.

I ran them locally on 2026-09-30 (macOS, Python 3.11, PyTorch 2.13, CPU) and hit two things:

  • The untouched starter code does not import. The body of QueryEmbedder.forward in dit.py is empty between BEGIN/END STUDENT SOLUTION, so Python raises IndentationError and test collection fails. I had to add a pass to all three TODOs before anything ran.
  • train_qformer.py does import wandb at the top level, so without wandb installed the tests don't even collect.

With the pass stubs, 13 tests fail and 1 passes (the submitted-files check). That is the expected starting state.

Training, inference, and the empirical questions

run_in_cloud.ipynb assumes Colab with the handout copied into your Google Drive. It mounts Drive, runs pip install -r requirements.txt and bash ./download_data.sh, then trains:

python train_qformer.py \
    --pretrained_model_path ./data/ddpm_dit_cifar_100_epochs.pth \
    --dense_captions_path "./data/cifar10_dense_captions.jsonl" \
    --epochs 25 \
    --batch_size 128 --lr 1e-4 \
    --save_model_path ./models/trained_qformer.pth \
    --gpt2_layer_index 12 --num_query_tokens 4 --cfg 3.0 --data_dir ./data \
    --gpt2_cache_dir ./data --optimizer_ckpt_interval 5 \
    --cache_text_embeddings ./data/text_embeddings.pt

Training pairs CIFAR-10 images with four kinds of caption: the class name alone, a synonym for the class name, a dense caption with a scene, and a synonym plus a scene. GPT-2 hidden states are cached to data/text_embeddings.pt. A footnote says the course could have had you generate them, but it would add time to an already long assignment. Text conditioning is dropped with probability 0.1 during training so classifier-free guidance works at inference time.

Inference is python eval_qformer.py --config_yaml inference.yaml. The inference.yaml in the zip has two examples: sampling from the original class-conditional DiT, and sampling "a photo of a silver plane in a green grassy background" from your trained Q-Former while saving attention maps at t = 1000, 500, and 100. You add your own entries for the other experiments.

The empirical questions (6.1–6.8):

ItemContentPoints
6.1Dimensions of GPT-2's hidden_states2
6.2Reading QueryEmbedder.__init__: the query table, each module's job, where Q/K/V come from, pre-norm12
6.3Count the Q-Former's trainable parameters and their share of the total; paste 3 dense-caption image-text pairs4
6.4Compare training only the Q-Former against unfreezing everything: asymptotic loss and the LLM's general ability4
6.5Train 25 epochs with num_queries=4; paste the EMA loss curve and sample grids5
6.6CFG scales w = 1, 2, 3, 4 on the original DiT for class "deer"3
6.78 cars each from class and text conditioning; two underspecified out-of-distribution prompts4
6.8Cross-attention maps of both layers at t = 500 for "a photo of a red airplane with a green field in the background"6

The hint for 6.8 points to Darcet et al. 2023 (Vision Transformers Need Registers) and Xiao et al. 2023 (attention sinks), and asks you to explain which tokens get attention and which are ignored completely.

Environment and compute: does it still run today?

Compute. Question 6.5 estimates 2–3 hours on a T4 and about 1 hour on an A100 for 25 epochs. The handout suggests claiming Colab Pro with a CMU email and choosing an A100 if available, with L4 or T4 as fallbacks. Outside readers don't get the education offer, so whether a free T4 survives 2–3 hours without a disconnect depends on your quota that day. The script saves weights and optimizer state every 5 epochs, and you can resume from a multiple of 5 with --resume_optimizer_path, --resume_model_path, and --start_epoch.

Downloads. download_data.sh uses gdown to fetch two files from Google Drive. On 2026-09-30:

  • cifar10_dense_captions.jsonl (about 7.6 MB) downloaded directly. Each line has index, split, label_index, label_name, and caption.
  • ddpm_dit_cifar_100_epochs.pth (Drive shows 51M) first returns a "can't scan for viruses" confirmation page because of its size. The file itself is still public.

Both files live on course staff's Drive, not on the course site. If the sharing changes, the assignment stops running, so self-learners should download a backup.

Dependencies. requirements.txt lists torch, wandb, Pillow, torchvision, tqdm, transformers==4.46.3, and gdown. Only transformers is pinned. I have 4.57.6 locally but did not run full training with it, so I can't say whether newer versions work. wandb is required: sample grids go to wandb, and 6.5 asks you to paste them from there.

Where the handout and the zip disagree:

  • The handout's file list omits test_all.py and inference.yaml, which are both in the zip. It also shows a handout/ folder, but the zip puts files at the root.
  • Question 6.5 says to train 25 epochs, then asks for sample grids from epochs 5, 25, and 49. The script's --epochs default is 20; the notebook command says 25.
  • Part 2 says the DiT has about 100M parameters and GPT-2 about 117M. Question 6.3 then says "suppose GPT-2 has 125M and DiT has 18M." Use the numbers in 6.3 when answering 6.3.
  • The points table gives the VLM section 18 points, but the per-item points add up to 11 (7 for 5.1, 2 for a second item also numbered 5.1, and 2 for 5.2). The total of 79 follows the table.

What you can't get

The March 13 HW4 recitation is a Google Slides link that returns 401 from outside CMU. Also out of reach: Gradescope autograding and grading rubrics, the Programming Test, official solutions, and the recordings (Panopto needs a CMU login).

What to do: tonight, download hw4.zip, add a pass to each of the three TODO bodies, install wandb, and run pytest test_all.py until you see 13 failures. Then open dit.py, read only QueryEmbedder.__init__, and write each module's input and output shapes next to it. The 12 points in 6.2 are mostly that shape table.

What this post can and can't confirm

Confirmed: the handout, starter code, unit tests, and notebook in hw4.zip; schedule dates; the reminder slides in L15 and L18; the Drive files' availability on 2026-09-30; and my local unit-test results. Not confirmed: real training time on a free T4, compatibility with newer transformers, the recitation content (401), whether the tentative Slot B date on the schedule (March 30) held that semester, and official solutions.

Further reading: for the math behind diffusion and guidance, see Reading MIT 6.S184. For the systems side of parameter-efficient fine-tuning such as LoRA, see CMU 11-868 L23: Efficient Fine-Tuning for Large Models.

Series: Previous: L14–L15: Cross-Attention, DiT, Prompt-to-Prompt, and the Q-Former | Next: L15–L16: Scaling Laws and Mixture of Experts | Series overview

References