Where does the text condition enter an image generator? CMU 10-423 L14 answers with cross-attention: queries come from the image's latent representation and keys and values come from the prompt, so every latent pixel gets a probability distribution over which words to look at. That attention map is useful. Classifier-free guidance makes generations follow the prompt more closely, and Prompt-to-Prompt copies old attention maps into a run with an edited prompt so only part of the image changes, with no retraining. DiT swaps the UNet for a Transformer and injects conditions with adaLN-Zero. In the first half of L15, the Q-Former uses a small set of learnable queries to connect a frozen image encoder to a frozen LLM, which is what HW4 asks you to build.
HW4 in CMU 10-423 Spring 2026 is worth 79 points. The written part covers LDMs (7), VQ-VAEs (8), CLIP (4), and VLMs through PaliGemma2 (18). The programming part (40) has you train only a Q-Former between a frozen GPT-2 and a frozen CIFAR-10 DiT, so a class-conditional diffusion model learns to take text. You write three functions, checked by 14 unit tests. The handout estimates 2–3 hours on a T4 or about 1 hour on an A100 for 25 epochs, and the captions and DiT weights come from Google Drive via download_data.sh.
Lab 3 builds a conditional latent diffusion model on MNIST from scratch, in four stages: CFG training with label dropout (checked on a three-component Gaussian mixture), a diffusion transformer built piece by piece (Fourier time embedding, patchify, multi-head attention, adaLN-Zero, depatchify), a VAE, and finally the DiT trained inside the VAE's latent space. Problems and official solutions are public; submission goes through Gradescope on Canvas, which only enrolled MIT students can use.
The algorithms are complete by Lecture 3B; Lecture 4 tackles two engineering problems that show up at scale. First, the network must take an image, a time t, and a prompt and output a vector field of the same size, so we use a U-Net or a diffusion transformer (DiT), embedding time with Fourier features and text with frozen CLIP/T5 encoders. Second, pixel space is too big, so we first train a VAE to compress images into a latent space, run flow matching there, and decode at the end. Stable Diffusion 3 and Meta Movie Gen Video both follow this recipe: flow matching in latent space, a DiT variant, and CFG.
Sora's February 2024 preview shocked the industry, but the landscape reversed in two and a half years: OpenAI closed the consumer Sora app in April 2026 and scheduled its API for retirement on September 24; Veo 3.1 became the narrative default with native audio and Flow; Kling 3.0 became a unified multimodal model with $240M annualized revenue; and Runway Gen-4.5 briefly led Artificial Analysis in a November 2025 snapshot while defending the professional market through enterprise workflows. This guide compares four families by generation, specifications, pricing, and use case.
Every serious image-to-video model in 2026 runs latent diffusion on a DiT backbone, so visual quality is no longer a useful axis for choosing one. The real axes are native audio, self-hostability, and dollars per second. Three widely-repeated errors worth correcting: Sora's app shut down on April 26 and its API goes on September 24; Wan 2.7 is described everywhere as Apache 2.0 open weights but no first-party source has them; Veo 3.1 officially costs $0.40/s, not the $0.75/s that circulates on review sites.