Skip to content

MIT 6.5940 Fall 2026 Lab 1 Supplement: Reading GPU Bottlenecks with Roofline, the Profiler, and FlashAttention

Sep 30, 20261 min
TL;DRThis post covers Fall 2026 material, not the Fall 2024 edition the rest of the series follows. Fall 2026 replaced the pruning lab with "Efficient AI Fundamentals" (lab1_gpu_basics.zip). Part 1 has you hand-write a triple-loop GEMM and compute MAC, FLOPs, and I/O. Part 2 plots GEMM and GEMV rooflines. Part 3 works through a gemma-3-270m-it decoder layer, computing attention and MLP costs and comparing prefill with decode. Part 4 uses the PyTorch Profiler to inspect kernels, has you write GeLU to feel kernel fusion, then tries torch.compile and CUDA Graphs. Part 5 compares SDPA with FlashAttention. The core is 80 points plus 20 bonus, and all of Part 5 became bonus because Colab's T4 can't run it.

🌏 中文版

This post covers Fall 2026 material. The Reading MIT 6.5940 series follows Fall 2024. This is post 17, and the only one built mainly on Fall 2026 material.

Why it's included: Fall 2024 has no GPU profiling lab, yet Lecture 13 keeps leaning on "decode is limited by bandwidth" and "FlashAttention moves less data". Fall 2026's Lab 1 lets you measure those claims yourself. It sits after Lab 4 + Lab 5 as the closing exercise for the LLM inference stretch.

Series position: previous Lab 4 + Lab 5: AWQ and LLaMA2-7B on a laptop | next L14 LLM post-training | series overview

Official materials: lab1_gpu_basics.zip, linked from the Lecture 4 row of the Fall 2026 course page. The files inside are dated 2026-09-29. The schedule has it released September 22 and due October 1 (Lecture 7). Question numbers and points below follow the README and notebook in the archive, checked on 2026-09-30.

Access level A2 (semester in progress): the archive is publicly downloadable, but submissions go through MIT's Canvas, there are no public solutions, and the course page says it isn't taking cross-registered students this semester. The notebook includes a few public test cases for self-checking. This post doesn't include solutions.

What's in the archive

The README's title is "Lab1: Efficient AI Fundamentals", and its acknowledgment credits Zhijian Liu's Z Lab for providing the lab.

  • lab1(colab).ipynb: for Colab, with images embedded. Upload the whole lab1_gpu_basics folder to Google Drive, open the notebook from there, and run Setup first.
  • lab1.ipynb: same content, for a local machine with an NVIDIA GPU.
  • utils/: benchmarking, a GPU spec table (config.py includes T4, L4, A100-40GB, A6000, H100, and others), and roofline plotting helpers.
  • Submit only one of the two notebooks.

Pin the versions: the README says pinned versions such as transformers==4.57.6 are required, because newer transformers releases changed an API that utils/config.py relies on. Colab's preinstalled transformers is too new, so skipping Setup fails at import.

GPU requirements: Parts 1–4 work on a Colab T4. Part 3.3's prefill/decode latency measurements need an A100 or A6000, but that section has no questions, so T4 users can skip it. For Part 5 on Colab, set the runtime version to 2025.10 and use an A100 so a prebuilt flash-attn installs. Otherwise pip compiles it from source, which can take hours.

The five parts and their points

The core is 80 points:

PartTopicQuestions (points)Subtotal
1Basic metrics: latency, MAC, I/O1.1 hand-written triple-loop GEMM (5), 1.2 latency measurement (5), 1.3 MAC for GEMM/GEMV (5), 1.4 I/O for GEMM/GEMV (5)20
2Roofline2.1.1 GEMM roofline (5), 2.1.2 how large N must be to go compute-bound (5), 2.2 GEMV roofline (5)15
3Gemma-3 decoder layer case study3.1 attention MAC (10), 3.2 attention I/O (5), 3.4.1 prefill vs decode (5), 3.4.2 effect of batch size (5), 3.4.3 effect of prefill length (5)30
4Profiling and kernel fusion4.3.1 write GeLU (3), 4.3.2 GeLU latency (2), 4.3.3 findings (2), 4.3.4 explain with the profiler (5), 4.4.1 compiled MLP (3)15

Bonus, 20 points: MAC and I/O for the whole attention module (2.5 each), profiling the decoder layer during decode (2.5), CUDA Graph profiling (2.5), FlashAttention vs SDPA (5), and FlashAttention analysis (5).

Parts 1–2: describing an operation with three numbers

Setup. Part 1 asks for a triple-loop matrix multiply in plain Python, with no NumPy or PyTorch ops. The notebook explains this is deliberate: the loop runs on the CPU, and later you compare it with torch.matmul.

Three metrics. Latency measures time. A MAC is one multiply-accumulate, and 1 MAC equals 2 FLOPs. I/O is how many bytes have to move. Question 1.2 asks you to time your GEMM several times in a row without warmup and describe what you see. It builds the instinct that measurement itself is noisy.

Roofline. Part 2 defines arithmetic intensity as FLOPs ÷ bytes and the ridge point as peak FLOPs ÷ peak bandwidth. In the memory-bound region on the left, performance = bandwidth × arithmetic intensity. In the compute-bound region on the right, performance = peak FLOPs.

2.1 plots FP32 GEMM rooflines for N from 1024 to 8192 and asks you to derive how large N must be on an A6000 before GEMM turns compute-bound. 2.2 repeats the plot for GEMV. Put the two plots side by side and you'll see what Lecture 13 means by "decode is GEMV, and it's slow because of moving weights".

Remember to change the notebook's default gpu_name = "A6000" to the card you're actually using. The comment lists the Colab options.

Part 3: working through a real decoder layer

The case-study model is gemma-3-270m-it. The notebook first lists every component of one decoder layer: input LayerNorm, self-attention (Q/K norm, QKV projections, rotary embedding, output projection), two LayerNorms, the MLP (gate, up, activation, down), and a post-MLP LayerNorm.

It demonstrates how to compute the MLP's MAC and I/O, then asks you to follow the same conventions for attention (3.1, 3.2). The notebook says any extra assumptions get full credit as long as they're reasonable and stated clearly in a comment.

Next it plots the MLP's prefill and decode on the roofline and states the conclusion itself: prefill is compute-bound, decode is memory-bound, and this isn't just an MLP quirk but a general feature of modern LLM inference. The three 3.4 questions dig into why:

  • 3.4.1: explain the difference from the query's shape (hint: think GEMM versus GEMV).
  • 3.4.2: during decode, raise the batch size from 1 to 64 and track how the point moves on the roofline.
  • 3.4.3: during prefill, raise the length from 64 to 4096 and track how the point moves.

The answer to 3.4.2 is the reason Lecture 13 says W8A8 suits batched serving while single-user decode needs W4A16.

Part 4: see the kernels, then merge them

Profiler. 4.1 uses torch.profiler to show which CUDA kernel a single torch.matmul actually calls. 4.2 switches to the MLP, where the kernel count jumps, and shows how to export a Chrome trace and open it in Perfetto as a timeline.

Kernel fusion. The notebook borrows Horace He's factory-and-warehouse analogy. A unary op like torch.cos ships data from the warehouse (memory) to the factory (compute units), does a tiny bit of work, and ships it back, so nearly all the time goes to shipping. A chain of such ops makes that round trip at every step. Fusion keeps the data in the factory until all the work is done.

4.3 lets you feel this directly. Write the tanh-approximated GeLU from basic PyTorch ops, compare its latency with torch.nn.functional.gelu(x, approximate="tanh"), and use the profiler to explain where the gap comes from.

torch.compile and CUDA Graphs. 4.4 first hands your GeLU to torch.compile to see what it fuses into, then compiles the MLP with max-autotune-no-cudagraphs and compares the before-and-after profiles. Bonus 4.4.2 switches to max-autotune (which enables CUDA Graphs), and the graphs only replay under torch.no_grad(). The notebook also warns that on GPUs with few SMs, like the T4, max-autotune doesn't autotune the matmuls, so the main visible change is the fused element-wise ops.

Part 5: FlashAttention (all bonus now)

The notebook says this part became bonus because a standard Colab T4 can't run it, and the staff didn't want students stuck hunting for GPUs near the deadline.

  • 5.1: measure standard SDPA's peak memory at sequence lengths from 512 to 32768, then answer what its memory complexity is, at what length it becomes a problem on your GPU, what its main memory overhead is, and which LLM scenario that would hurt.
  • 5.2: compare the latency and memory of flash_attn_func against SDPA, then use the profiler to see how FlashAttention fuses attention into a single kernel.

The notebook's explanation of FlashAttention matches pages 82–83 of Lecture 13: it never writes the $N \times N$ attention matrix to HBM, splits Q, K, and V into blocks that fit in SRAM, and computes block by block with an online softmax. It lists the payoff as memory dropping from $O(N^2)$ to $O(N)$, a 2–4x speedup from fewer HBM accesses, and exact rather than approximate attention.

How this lab maps to the Fall 2024 spine

What this lab practicesWhere Fall 2024 covers it
Definitions of MAC, FLOPs, latencyLecture 2: efficiency metrics
Prefill and decode have different bottlenecksLecture 13, pages 19–20
Kernel fusionLecture 13, page 37 (TinyChat fusing dequantization with the matrix multiply)
FlashAttentionLecture 13, pages 82–83
Parallelism and the hardware underneathLecture 11: TinyEngine

The difference is that Fall 2024's Labs 4 and 5 have you optimize kernels on a CPU, while this lab has you measure and diagnose on a GPU. Together they make the full loop: find the bottleneck first, then optimize.

Also note that because Fall 2026 swapped in this lab, it no longer has a pruning lab. To practice pruning, use Fall 2024's Lab 1.

How to self-study it

  1. Decide where you'll run it. With only Colab's free T4, do Parts 1–4 and skip Part 3.3 and Part 5. With an A100 or A6000, do the full version.
  2. Restart the runtime after Setup. The README calls this out: skip the restart and imports fail.
  3. Work Parts 1–2 on paper first. Derive the MAC and I/O for GEMM and GEMV by hand, then write the code and check it against the public test cases.
  4. In Part 4, export a trace at least once and open it. The table shows only totals. The timeline shows the gaps between kernels.

One thing you can do tonight: download the archive and do only 2.1.2, using your own GPU's peak compute and bandwidth to find its ridge point. After that, you can check any "this op is memory-bound" claim yourself.

Further reading

References