Skip to content

CMU 11-868 L06–L07: Reading Transformers, T5, LLaMA, and GPT-3 Like a Systems Engineer

Sep 30, 20261 min
TL;DR11-868 spends only two lectures on the model itself. L06 breaks the Transformer into embeddings, multi-head attention, FFN, LayerNorm, and residuals; L07 uses T5, LLaMA, and GPT-3 to show what modern LLMs changed. For a systems engineer the point is to remember the shapes: GPT-3 175B has 96 layers, d_model 12288, a 2048-token context, and trained on 300B tokens; LLaMA 65B has 80 layers, d_model 8192, and trained on 1.4T tokens. Those numbers set the workload for every acceleration, parallelism, and serving lecture that follows.

🌏 中文版

Version note: This post follows the Spring 2026 offering of CMU 11-868 LLM Systems. The main sources are the L06 Transformer slides (Feb 2, 25 pages), the L07 Pre-trained LLMs slides (Feb 4, 22 pages), and the readings listed in the Syllabus. Page numbers refer to PDF pages. All facts were checked against the official materials on 2026-09-30. Access level A3: slides and assignments are all public, but there are no public recordings, so this post works from slides and papers only and cannot relay anything the lecturer said aloud.

Series: previous HW2: MiniTorch Framework | next L08–L09: Tokenization, decoding, and speculative decoding | Series overview

The first five posts laid the foundation: how a GPU runs a kernel, how a framework differentiates automatically, and how your own MiniTorch trains a sentiment classifier. This is where the course puts an actual model on the table for the first time.

11-868 covers the model in just two lectures: L06 on the Transformer, L07 on pre-trained LLMs. Compared with Stanford CS224N or CME295, that is short. The goal here is not to explain why language models work. It is to show you what the thing you will accelerate, shard, and serve actually looks like. This post reads the lectures from that angle and asks two questions of every component: what shape are its matrices, and what resources does it consume?

Orientation: three kinds of language models

L06 page 4 sorts language models into three types:

TypeTraining objectiveExamples on the slide
Encoder-onlyMasked LMBERT, RoBERTa, ESM (protein)
Encoder-decoderAutoregressive or non-autoregressiveT5
Decoder-onlyAutoregressiveGPT, LLaMA, ProGen (protein)

L06 uses machine translation as its running example and follows the encoder-decoder route; L07 switches to decoder-only. That matters for the homework: HW3 asks you to do German-English translation with a decoder-only GPT-2 architecture, so you need both.

L06 pages 6–7 give the motivation for the new architecture. Earlier seq2seq models used LSTMs or GRUs and processed one position at a time. The Transformer uses attention in both encoder and decoder; with recurrence gone, the encoder can encode the whole sentence at once. For a systems course, that is the whole point: parallel work is what GPUs are good at.

L06: the shape of one Transformer layer

Embeddings

L06 page 10: the token embedding is a lookup table shared (tied) between input and output. Positional encoding uses the original paper's sin/cos formula, has the same dimension as the token embedding, and is added to it. Tokenization is deferred to the next lecture.

The systems takeaway: the embedding table is vocabulary size × hidden dimension. A bigger vocabulary means a bigger table and a bigger output layer, a cost the next post comes back to.

Multi-head attention

L06 pages 11–13: the input X is (number of tokens × dimension). It is projected into Q, K, and V, split into h heads, each head runs scaled dot-product attention, and the results are concatenated and multiplied by the output matrix W^O. Page 12 labels two shapes on the diagram: Q, K, and V are len × dim, and the attention score matrix is len × len. It also leaves a question for the reader: why divide by √d?

That len × len is the setup for the rest of the course. Double the sequence length and the attention matrix gets four times bigger. The later FlashAttention lecture is entirely about not writing that matrix back to memory in one piece.

Decoder self-attention adds one step: positions to the right (the future) are masked to −∞ before the softmax (page 14).

FFN, residuals, and LayerNorm

The FFN on L06 page 13 is two linear layers with a ReLU: FFN(x) = max(0, xW1 + b1)W2 + b2. Page 17 gives the original paper's numbers: 6 encoder and 6 decoder layers, embedding size 512 (base) or 1024 (large), FFN hidden size 2048.

Page 15 puts residual connections and LayerNorm together and lists two placements, post-norm and pre-norm. In HW3 this becomes a requirement: the assignment page states that GPT-2 uses pre-LN.

Training setup

L06 pages 18–23 summarize the original paper's training details. The ones that bear on resources:

  • Training uses teacher forcing: the decoder pretends it already knows the correct prefix, so loss for every position is computed together (page 18)
  • Batches are grouped by approximate sentence length but still shuffled (page 21)
  • Hardware: the 2017 paper used one machine with 8 GPUs; the base model took 100k steps (about 12 hours) and the large model 300k steps (about 3.5 days) (page 21)
  • Adam with a learning rate that warms up and then decays (pages 21–22)
  • The base model averages its last 5 checkpoints (page 23)

The last page of L06 (page 25) points the code walkthrough to The Annotated Transformer, and the Syllabus schedules it as Recitation 3 on Feb 6. Without recordings, this line-by-line implementation is the best substitute for L06.

L07: what modern pre-trained LLMs changed

L07 uses three models as case studies (page 3): the encoder-decoder T5, the decoder-only LLaMA, and GPT-3.

T5: a unified format and span corruption

Pages 4–7:

  • A standard encoder-decoder Transformer, decoded with beam search (beam width 4, length penalty 0.6)
  • Sizes: T5-base has 220M parameters (12 blocks, d_model 768, d_ff 3072, 12 heads); T5-11B has 24 blocks, d_model 1024, d_ff 65536, and 128 heads
  • Pre-training data is C4, a filtered English corpus from Common Crawl, 750GB
  • Pre-training objective: corrupt 15% of the text as random spans and recover them; 0.5M steps with batches of 128 sequences of length 512, packed to about 65k tokens per batch, about 34B tokens in total
  • Multitask fine-tuning writes the task instruction into the input as natural language; T0 and Flan-T5 build on this

T5-11B's d_model is only 1024; almost all of its parameters sit in the 65536-wide FFN. It is a reminder that the same parameter count spread over different matrices puts very different pressure on memory and compute.

LLaMA: three architectural changes

Page 8 lists three changes LLaMA makes to a decoder-only Transformer, each with its source:

  1. Pre-normalization (from GPT-3): LayerNorm moves before each sublayer (page 9)
  2. SwiGLU (from PaLM): the FFN uses a Swish gate; page 10 notes the hidden size changes from 4d to 2/3 × 4d
  3. RoPE (from RoFormer): a rotation matrix makes attention scores depend only on relative position (pages 11–12)

Page 14 adds the training recipe: standard language-modeling loss without label smoothing, plus an auxiliary loss that keeps the softmax normalizer close to 0; pre-training uses only open-source data.

The model-size table on page 13 is an image, so the numbers come from Table 2 of the LLaMA paper:

Paramsd_modelHeadsLayersTraining tokens
6.7B409632321.0T
13.0B512040401.0T
32.5B665652601.4T
65.2B819264801.4T

Section 2.4 of the paper has a number a systems course will like: training the 65B model processed about 380 tokens per second per GPU on 2048 A100s with 80GB, so one pass over 1.4T tokens took roughly 21 days.

GPT-3: size and compute

Page 15: GPT-3 keeps the standard Transformer but changes the initialization, uses pre-normalization and reversible tokenization, and alternates dense and locally banded sparse attention. The size table, training details, and "Computation" slide on pages 16–18 are screenshots from the GPT-3 paper. Table 2.1 of the paper gives the largest model, 175B, as:

  • 96 layers, d_model 12288, 96 heads of 128 dimensions each
  • A 2048-token context window for every model size
  • 300B training tokens for every model size

Table D.1 in Appendix D does the compute accounting: about 2 floating-point operations per parameter per token in the forward pass, times 3 to include the backward pass, for 6 in total. Training 175B on 300B tokens comes to about 3.14 × 10²³ FLOPs, or about 3,640 petaflop/s-days.

Back-of-the-envelope math

The following are my own rough estimates from the numbers above, not slide content. But this is exactly the exercise the rest of the course asks of you:

  • Parameter memory: 175B parameters in FP16 is about 350GB for the weights alone, more than one 80GB GPU holds. That is why model parallelism and ZeRO exist
  • Parameters per layer: attention (four d×d matrices for Q, K, V, O) plus the FFN (two d×4d matrices) is about 12d². With d = 12288 and 96 layers, 12 × 12288² × 96 ≈ 174B, which matches 175B
  • Attention scores: with a 2048 context, each head in each layer has a 2048 × 2048 score matrix

Further reading: the architecture itself

11-868 deliberately covers only what systems work needs. For the reasoning behind the architecture, the site has fuller guides:

How to self-study this part

  1. Read L06, then work through The Annotated Transformer and write every matrix shape down on paper, especially the permutes and reshapes in attention. HW3 tests exactly that
  2. L07's size tables are images; open Table 2 of LLaMA and Tables 2.1 and D.1 of GPT-3 directly
  3. Use GPT-3's numbers to compute parameter count, weight memory, and training FLOPs once, as a warm-up for the second half of the course
  4. L07 page 21 links straight to the HW3 assignment page. The assignment went out on Feb 4; after this post and the next one you can start it

References