Skip to content

Reading Stanford CS224U, Part 17: Two Extension Lectures — Generating Text with Diffusion, and Turning LLM Training Folklore into Intuition

Sep 29, 20261 min
TL;DRIn Spring 2023, CS224U slipped two talks by members of its own teaching team into the Transformer unit. Lisa Li presented Diffusion-LM: instead of generating left to right one word at a time, it denoises a sequence of Gaussian vectors into word vectors. It loses to autoregressive models on both training and decoding efficiency, and in exchange lets a classifier's gradient steer the output at every step. Sidd Karamcheti showed how to cut GPT-2 Small's single-GPU training clock from 99.63 days to 3.37 days by stacking data parallelism, mixed precision, and ZeRO. The first talk survives only as slides; the second has slides and two recordings.

🌏 中文版

This is the last post in the Reading Stanford CS224U series, and the only one that's optional.

The Spring 2023 CS224U course site lists four items for its first unit starting Apr 5 (Domain adaptation for supervised sentiment): the Assignment 1 overview, Contextual word representations, and two decks presented by someone else — Diffusion objectives for text (Lisa) and Fantastic language models and how to build them (Sidd). Both speakers, Xiang (Lisa) Li and Sidd Karamcheti, appear on the site's Teaching team list, so strictly speaking these aren't outside guest lectures; they're special topics led by team members.

This series moves both talks from unit one to the very end, for a simple reason. Part 4 has just finished the GPT, BERT, and ELECTRA model families; dropping "text can be generated by diffusion too" and a full distributed-training toolkit right after that means two big steps on the same stretch of road. Neither talk connects directly to the assignments or the final project. They're a slice of what this group was thinking about in 2023, outside the main line of the course.

What you can get

LectureOfficial materialRecordingAccess
Diffusion objectives for text (Xiang Lisa Li)88-page deck; the schedule lists Diffusion-LM (Li et al. 2022) as a readingNot among the 50 videos in the XCS224U playlistSlides and paper only
Fantastic Language Models and How to Build Them (Siddharth Karamcheti)27-page deck, cover dated April 12, 2023Playlist videos 49 and 50Slides and recordings both public

The series as a whole rates A3 (on the scale from the global AI/CS course map: materials and assignments sufficient for self-study). This post on its own is closer to A2: slides and some recordings, no exercises. The speaker's narration for the diffusion talk is gone entirely, and the result bar charts in that deck carry no numbers, so wherever results come up below, I can only say which methods were compared.

Lecture one: can we generate every word at once?

The deck opens with a row of names. DALL·E 2, Imagen, and DDPM use diffusion for images and audio; GPT-2, GPT-3, PaLM, OPT, GPT-4, Claude, and LLaMA are all autoregressive. The slide title is "Monopoly of Autoregressive LMs."

Two limits of autoregressive models

Li demonstrates an autoregressive LM with "Harry Potter graduated from ___": the model assigns probabilities to the next word (Hogwarts 0.8, Oxford 0.05), samples one, appends it, and computes the next. A sentence's probability factors via the chain rule into a product of per-position conditionals.

From there the deck names two limits:

  1. Time complexity is O(n), where n is the length of the text; words come out one at a time.
  2. The generation order is fixed. Generating right to left, or filling in the middle given left and right context, doesn't fit the architecture naturally.

Then comes the question for the whole talk: "Can we generate all words at once?" The answer is Diffusion-LM, a NeurIPS 2022 paper by Li with John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori Hashimoto.

How image diffusion works

A few slides cover image diffusion quickly. Generation starts from pure Gaussian noise x_T; the model learns p(x_{t−1} | x_t) at each step and denoises its way back to a clean image x_0. Training runs the other way: add noise to real data step by step to build (x_t, x_{t−1}) latent pairs, then train supervised so the model's predicted mean matches the true posterior mean.

The deck sums up the process in one line: "Construct latent variables pairs, then apply supervised training."

Two problems when moving to text

Images are continuous; text is discrete words. Diffusion-LM has to deal with two things.

First, words have to become vectors. Each word is mapped to a d-dimensional vector (d is a hyperparameter, described as a small number), so a sentence becomes a continuous n×d matrix, and diffusion runs in that space. Where do the embeddings come from? The deck weighs two options, random embeddings or learning them end to end. Diffusion-LM learns them end to end, adding a reconstruction term (the probability of recovering the original words from x_0) to the denoising loss so the embeddings train alongside the model.

Second, the last step has to round. Once denoising reaches x_0, each position takes its most likely word. Ideally x_0 lands exactly on a word embedding; in practice it doesn't. The deck's example sticks: in "My ice cream is [BLANK]," melting is right and saving is wrong, yet the two words sit close together in embedding space.

Li's fix is a change of training target. Originally each step predicts the previous time step, which pushes all the rounding into the final x_1 → x_0 step, where it's hard and error-prone. Instead, every step predicts the clean x_0 directly. The model is forced to align with real word embeddings at every noise level, the predicted x_0 gets more precise, and rounding error drops.

Where text differs from images

Two slides discuss frequency. In images, low frequencies carry large-scale structure, medium frequencies carry detail, and high frequencies are perceptually almost meaningless, so adding a bit of high-frequency noise leaves an image looking about the same.

What does "high frequency" even mean for text? The deck offers one definition: the rate of change between neighboring word embeddings. The catch is that this component matters a lot for text. It bears directly on token prediction, and readers spot a bad n-gram instantly. What images can ignore, text can't.

Against autoregressive models, on four axes

AxisThe deck's verdictWhy
Training efficiencyAutoregressive winsWith causal masking, one forward pass gives an autoregressive model gradient signal at every position; one forward pass of Diffusion-LM trains only one noise level
Decoding efficiencyAutoregressive winsAutoregressive is O(sequence length), Diffusion-LM is O(number of diffusion steps); but autoregressive models can cache, while diffusion recomputes everything each step. The deck adds that Diffusion-LM may pay off more for longer text
Generation orderDiffusion-LM winsNot tied to left to right
Controllable generationDiffusion-LM winsSee the next section

Along the way there's a comparison of writing styles: humans write long text as core concepts → structure → wording; autoregressive models go left to right; Diffusion-LM goes coarse to fine.

The real selling point: controllable generation

The last section is what the paper is actually about. Say the goal is "write a positive review of Coupa, a coffee shop on the Stanford campus," and all you have is a pretrained model with frozen parameters.

The plug-and-play approach brings in a separate critic p(c | x) and uses Bayes' rule, p(x | c) ∝ p(x) p(c | x), to pull generations toward the target. The deck walks through repeated attempts against an autoregressive model: it produces the three little pigs, then Starbucks, and the critic keeps rejecting.

Diffusion-LM's advantage is its chain of continuous intermediate states x_T, …, x_0. At each denoising step, besides following the model's own transition p(x_{t−1} | x_t), you can also update x_{t−1} along the gradient of the classifier score p(c | x_{t−1}). The control signal gets to act at every noise level instead of scoring the sentence only after it's finished.

The experiment slides show "controlling semantic content": given a field and value (e.g., Food = Japanese), generate a sentence that covers the value, scored by exact-match success rate, with fluency measured separately. The baselines are PPLM, FUDGE, and FT-sample. The bar charts have no numbers; the paper's abstract says Diffusion-LM significantly outperforms prior work on six fine-grained control tasks.

Lecture two: turning folk knowledge into intuition

Karamcheti's deck carries the subtitle "Stanford || Zoom || Folks 2x-ing the Recording." In the recording he introduces himself as a fourth-year PhD student working on language for robotics. His opening claim: every new GPT generation adds more "folk knowledge" that stays hidden and never gets written up, and the job of academics is to rediscover it and turn it into intuition.

The 27 pages come in three parts:

  1. The evolution of the Transformer. Starting from RNNs (long context, attention) and CNNs (multiple filters, residuals, parallelizable), it builds up to self-attention and multi-head attention, then keeps asking what's missing. No nonlinearity, so add an MLP (in the recording he uses the SVM kernel-lifting analogy). Activations blow up, so add LayerNorm. Then optimization problems appear, which leads to learning-rate warmup (the slide says "linear warmup for 5% of training, then decay" and notes it seems to break conventional ML wisdom). The final slide of the part is "The Modern Transformer (March 2023)."
  2. Training at scale. See the next section.
  3. Fine-tuning and inference. The tools learned for training carry straight over: ZeRO Infinity for CPU/NVMe offloading, 8-bit quantization (the slide cites LLM.int8() and says "Powers llama.cpp and more!"), and a teaser for parameter-efficient fine-tuning such as LoRA, pointing to Hugging Face PEFT.

How 99.63 days becomes 3.37

Part two is the most concrete stretch of the talk. Karamcheti starts from his own experience: he wanted to train GPT-2 Small (124M parameters), hit OOM with batch size above 4 on a 12 GB GPU, fixed that with gradient accumulation, and then found that 400K steps on a single GPU would take 99.63 days.

The stated goal is "100 days on 1 GPU → about 4 days on 16 GPUs," and the training clock updates with each technique added:

TechniqueTraining clock on 16 GPUsThe deck's point
Single GPU (baseline)99.63 days—
Distributed data parallel (DDP)7.2 daysA simple wrapper around nn.Module that auto-partitions data across processes
DDP + FP16 mixed precision6.01 daysStatic memory drops from 20 to 16 bytes per parameter; the real speedup comes from NVIDIA Tensor Cores
DDP + FP16 + ZeRO3.37 daysOptimizer states and gradients are sharded by GPU count instead of replicated on every card

The memory slide is worth a pause. With FP32 and Adam, each parameter needs the parameter, a parameter copy, the gradient, momentum, and variance, for a lower bound of "number of parameters × 20 bytes." From that the deck estimates 18 GB for 1B parameters and 3 TB for 175B, before activations. FP16 doesn't mean everything is 16-bit; optimizer states stay 32-bit, which is why it only gets down to 16 bytes.

The last slide admits the wall: communication between nodes eventually costs too much, and the next steps (splitting matrix multiplies, scheduling the backward pass) are "harder to implement, model-specific… still miles to go."

Which part of the recordings to watch

Neither playlist video is the whole talk. Video 49 (about 46 minutes) opens with Potts finishing the remaining contextual-representations sections, ELECTRA and after; Karamcheti takes over roughly two-thirds of the way in and stops at the MLP-and-kernels analogy when class ends. Video 50 (about 81 minutes) spends its first half on Potts wrapping up neural IR; Karamcheti gets the second half, from warmup through ZeRO.

Part three barely appears in the recordings. Time ran out, and he recommends PEFT in a single closing sentence. For the fine-tuning and inference content, slides 25–26 are all there is.

The unit reading: The Pile

The reading list for the same unit includes The Pile (Gao et al. 2020). The schedule lists readings for the unit as a whole without saying which talk it belongs to, and neither deck's text mentions it.

The paper presents an 825 GiB English corpus built from 22 subsets, many from academic or professional sources. Evaluating GPT-2 and GPT-3 on it, the authors find they struggle on components such as academic writing; models trained on the Pile improve significantly over Raw CC and CC-100 across all its components. The paper also documents potentially concerning aspects of the data and releases the construction code.

What you can do tonight

  • To understand Diffusion-LM: read slides 43–48 of the diffusion deck first (the rounding problem and "predict x_0 at every step"), then the paper's abstract and method section. Those pages carry most of the deck's information.
  • To understand training engineering: open slides 20–22 of the Fantastic LMs deck, multiply the parameter count of the model you use most by 20 bytes to get the FP32 + Adam static-memory floor, and see how many times over your GPU memory it is.
  • If you're watching the recordings, jump straight to the last third of video 49 and the second half of video 50.

Where it sits in the series

The previous post, Part 16: Writing NLP Papers, Submitting, and Giving Talks, is the end of the course's main line. This one is an extension beyond it, and the series ends here. Go back to the series overview for the map of the whole course, its offering status, and the assignment-environment pitfalls.

Further reading

References

All sources below were opened and checked on 2026-09-29: