Skip to content

MIT 6.5940 L12 Transformer and LLM: The Architecture Seen Through an Efficiency Lens

Sep 30, 20261 min
TL;DRIn Lecture 12, 6.5940 switches from CNNs to Transformers. The lecture doesn't dwell on theory. It points to where memory and compute go. Attention is O(N²). If Llama-2-70B used MHA, its KV cache at batch 16 and length 4096 would take 160GB. GQA shrinks that 8x and MQA shrinks it 64x. MoE adds total parameters while keeping per-token compute flat. This post bridges into Lecture 13 on LLM deployment.

🌏 中文版

Version note: This post is based on Lecture 12 (2024-10-17) of MIT 6.5940 Fall 2024. The main materials are Lec12-Transformers-and-LLM.pdf (90 pages) and the lecture recording. Page numbers refer to PDF pages. Facts were checked against the official materials on 2026-09-30. Access level A3: slides and video are public. This lecture has no matching lab.

Fall 2026 comparison: The F26 schedule puts the same lecture on October 22. As of 2026-09-30 its slide and video links are still empty.

Series: previous L11 TinyEngine and Parallel Computing | next L13 LLM Deployment | Series overview

Up to Lecture 11, nearly every example in 6.5940 is a CNN: pruning VGG, quantizing ResNet, squeezing MobileNetV2 onto an MCU. The F24 schedule inserts a "Chapter II: Domain-Specific Optimization" divider before Lecture 12. Everything after it (LLM deployment, post-training, long context, ViT, diffusion) is built on the Transformer. Lecture 12's job is to lay out that architecture so the later lectures share a vocabulary.

This post takes only the efficiency view: which design choices cost memory, which cost compute, and what each later variant saves. For the full theory, the site has better starting points: CS224N Lecture 5: from recurrence to Transformers, CME295 Lecture 1 and Lecture 2, and CS336 on architectures and hyperparameters.

The Lecture Plan on page 6 has four parts: Transformer basics, design variants, LLMs, and advanced topics (multimodal LLMs).

Transformer basics: where the cost hides

Why move away from RNNs and CNNs

Pages 9–13 first review the two approaches before the Transformer. RNNs struggle with long-range relationships: for two tokens to interact, information has to travel O(seq_len) steps. And since state n depends on the previous n−1 states, training is hard to parallelize. CNNs have no dependency between tokens and scale well, but their context is limited and their modeling power weaker. The slides also point out a difference between domains: images have locality, language doesn't necessarily.

What each component means for efficiency

Page 14 lists the Transformer's parts: tokenizer, embedding, Multi-Head Attention (MHA), Feed-Forward Network (FFN), LayerNorm, residual connections, positional encoding, and the final linear head. For efficiency, a few points are worth noting:

  • Tokenizer (page 16): one word can split into several tokens. The slide's example turns 110 words into 162 tokens. When we compute KV cache or attention cost later, the length N is the token count.
  • Self-attention (page 21): multiplying Q by K gives an N×N attention weight matrix, so attention compute is O(N²). The slides explain Q/K/V with a YouTube search analogy: the query is the text in the search box, the keys are video titles and descriptions, and the values are the videos themselves.
  • FFN (page 27): attention handles relationships between tokens but has no elementwise nonlinearity, so a two-layer MLP follows, with the hidden layer expanded to 4d. The slides call it an inverted bottleneck. Lecture 11's MobileNetV2 inverted residual block, which expands channels 6x, has the same shape.
  • LayerNorm and pre-norm (pages 29–30): Transformers use LayerNorm rather than the BatchNorm common in CNNs, normalizing each token's embedding separately. The slides note that pre-norm now beats post-norm in popularity because it trains more stably.
  • Positional encoding (pages 31–32): attention and the FFN don't distinguish order on their own. They treat a sentence as a set, so position information has to be added. The original design used absolute positional encoding.

The result on page 34: the original Transformer beat earlier models on machine translation at a fraction of the training cost.

Design variants: each one saves something

Page 36 lists four main groups of variants after the original paper:

  1. Encoder-decoder (T5), encoder-only (BERT), decoder-only (GPT)
  2. Absolute positional encoding → relative positional encoding
  3. KV cache optimization: MHA → MQA → GQA
  4. FFN → GLU

Three architecture families

Pages 38–41 cover these briefly. T5 casts every NLP task as text-to-text: the prompt goes into the encoder and the decoder generates the answer. BERT is encoder-only, pre-trained with a masked language model objective (randomly masking 15% of tokens) and next sentence prediction. GPT is decoder-only, pre-trained to predict the next word. Smaller models (GPT-2) get fine-tuned afterward; larger ones can work zero-shot or few-shot.

Relative positional encoding: train short, test long

Page 43 compares the two approaches. Absolute positional encoding fuses position into the input embedding, so it affects Q, K, and V and propagates through the whole network. Relative positional encoding affects only the attention scores (by adding a bias or modifying Q and K) and leaves V alone. The upside is possible generalization to lengths not seen in training: "train short, test long." The slide adds that this doesn't always hold.

  • ALiBi (page 44): adds an offset based on relative distance directly to the attention matrix, instead of to the token embeddings.
  • RoPE (pages 45–46): the approach LLaMA uses. Split the d-dimensional embedding into d/2 pairs, treat each pair as a 2D coordinate, and rotate it according to position m. The phase difference in the inner product of two vectors is then m−n, which depends only on relative position.

RoPE's efficiency payoff is on page 47. LLMs have a length limit set in training (LLaMA 2k, Llama 2 4k, GPT-4 8k) and break past it. Applying position interpolation to RoPE (using a smaller θ) extends LLaMA's context from 2k to 32k. This is the starting point for Lecture 15 on long context.

KV cache: the first thing to blow up at long context

Pages 49–51. When a GPT-style model generates token by token, each step needs only the current token's query, but it has to attend over the keys and values of every previous token. Storing those K and V tensors for reuse is the KV cache.

Page 52 gives the formula:

KV cache size = batch size × layers × kv-heads × head dimension × length N × 2 (K and V) × 2 bytes (FP16)

Plugging in real models:

ModelKV cache per token, per batch item
Llama-2-7B (32 layers × 32 heads × 128)512KB
Llama-2-13B (40 layers × 40 heads × 128)800KB
Llama-2-70B, assuming MHA (80 layers × 64 heads × 128)2.5MB

Page 53 scales up the 70B figure. Batch 1 at length 512 needs 1.25GB. Length 4096 needs 10GB. Batch 16 at length 4096 needs 160GB, which takes two A100s. The chart on page 54 shows that at length 2048, the KV cache quickly outgrows the model weights themselves as batch size increases.

This table is the page of this lecture that matters most for what follows. KV cache quantization, H2O, and PagedAttention in Lecture 13, and StreamingLLM and DuoAttention in Lecture 15, all attack this number.

MQA and GQA: store fewer heads

Pages 55–57 fix it by reducing the number of kv-heads:

  • MHA: N query heads, N key/value heads.
  • MQA: N query heads, a single key/value head.
  • GQA: N query heads, G key/value heads. The slide says G is typically N/8.

At Llama-2-70B's size (64 heads), GQA with 8 kv-heads makes the KV cache 8x smaller, and MQA with 1 makes it 64x smaller. Page 57 cites an experiment from the Llama 2 paper: at large enough model sizes, GQA matches MHA's accuracy. Page 70 also notes that Llama 2's 70B model uses exactly 64 heads and 8 kv-heads.

GLU: replacing the FFN

Pages 59–60: swapping the original FFN for a GLU variant (such as SwiGLU, a Swish activation plus an elementwise multiplicative gate) improves Transformer perplexity.

LLMs: design choices at scale

Pages 62–63 explain that an LLM is a Transformer scaled up and trained on a large corpus (natural language, code, and more). The slides plot model size against GPU memory on the same chart. Models grow much faster than single-GPU memory, which is one reason this course exists. Scaling also brings "emergent" abilities: some tasks only work once the model is big enough.

Pages 64–66 use GPT-3 (175B) to show in-context learning. With no fine-tuning, the model handles new tasks from a task description (zero-shot) or a few demonstrations (few-shot), and larger models make better use of the demonstrations.

Pages 68–72 list the design choices of several open models. For an efficiency reader, the takeaway is that most of these choices have converged:

ModelDesign choices listed on the slides
OPTdecoder-only, pre-norm (the 350M model is the exception with post-norm), ReLU in the FFN; nine sizes from 125M to 175B
LLaMAdecoder-only, pre-norm, SwiGLU, RoPE; 7B to 65B; context 2048
Llama 2context 2k → 4k; training tokens up from 1T/1.4T to 2T; GQA in the larger models; plus Llama-2-chat
Llama 315.6T training tokens; 50x Llama 2's compute; 405B flagship; 8K context for base, 128K for instruct; post-training with SFT, rejection sampling, and DPO
Mistral-7BGQA (8 kv-heads), 8k context, sliding window attention (dropped from v2 on); the 7B model beats Llama-2-13B

Chinchilla and inference cost

Page 73 cites Chinchilla: for a fixed training compute budget, scale model size and data together to get the best compute vs. accuracy trade-off. Then the slide adds a note that only an efficiency course would make:

Note: the trade-off is different if we consider the inference computation trade-off

If you care about inference cost, train a smaller model for longer, as LLaMA did. The slide points out that Llama-2 7B was trained on 2T tokens, far more than Chinchilla recommends. For deployment this is good news: a model of the same capability can be smaller.

Advanced topics: multimodality and MoE

Two ways to let an LLM see

Page 76 splits vision-language models into two groups:

  • Inject visual information with cross-attention (Flamingo style, pages 77–81): freeze the LLM and insert cross-attention layers between its layers. A Perceiver Resampler uses a small set of learned queries to compress image features of varying size into a fixed number of visual tokens. Gated cross-attention uses a tanh gate to control how much visual information flows in. The gate starts at 0, so at first it doesn't disturb the original LLM.
  • Feed visual tokens directly as input (PaLM-E style, pages 82–83): images, robot states, and other modalities are all turned into tokens for the LLM. RT-2 goes further and outputs control signals directly.

The Perceiver Resampler is itself an efficiency choice: the fewer visual tokens, the shorter the sequence the LLM has to process. This thread continues with multimodal LLMs in Lecture 14 on post-training.

MoE: more parameters, same compute per token

Pages 86–89 (citing Switch Transformers): MoE has each token use only part of the parameters. A router assigns tokens to different experts. More experts means more total parameters and lower loss, while per-token inference cost stays the same.

Page 88 explains the capacity factor C with a small example: 6 tokens, 3 experts. With C = 1, each expert handles at most 2 tokens, so one token gets skipped. With C = 1.5 the cap rises to 3, leaving experts 2 and 3 some slack. This parameter decides whether an unbalanced load drops tokens or reserves extra room.

What to do after this lecture

  • Tonight: pick an open model you use and read the layer count, kv-head count, and head dimension from its config. Use the formula on page 52 to compute the KV cache at batch 1 and length 8192, then compare it with the size of the weights. That ratio tells you whether your bottleneck is the weights or the KV cache.
  • Then read Lecture 13 on LLM deployment to see how quantization, sparsity, and serving each deal with these numbers.

Further reading

References