🌏 中文版
This post is based on the Fall 2025 edition of CS149. It is part 18 of Reading Stanford CS149. It follows L13 DSLs and AI-driven optimization and covers Programming Assignment 5 (stanford-cs149/asst5-kernels). The course home page lists it as "Assignment 5: Make the World's Fastest CUDA Kernels," due December 4, 2025. The README states there are no late days for it.
This post covers what each problem exercises, what to measure first, and which direction to think in. No solutions.
How this assignment differs from the first four
The README frames PA5 as "a very short final project." The staff calibrated it so a team can earn a decent score in about two evenings, while teams that want to go deep can spend much longer chasing very fast code. Compared with PA1 through PA4, three things change:
- No single task. Pick one or more of five AI-related kernels. Each comes with a PyTorch baseline.
- No performance bar. Your grade comes from the staff's reading of your work log.
- LLMs are allowed. The README says you may use any model in Stanford's AI Playground to write code, interpret profiler output, or decide what to try next.
The stated goal is practice in open-ended performance engineering, the real-world situation where you need a program to run faster and there is no fast staff reference to chase.
The target is an H100. Teams with PA4 credits left may optimize on Trainium instead, but Trainium has no leaderboard (see the series' PA4 Trainium2 + NKI post for that environment).
Three ways to write it
| Option | File you hand in | Notes |
|---|---|---|
| Python + Triton | submission.py, implementing the custom_kernel interface | The job queue runs triton 3.5.1 |
| Python + TileLang | submission.py | The job queue runs tilelang 0.1.6.post2 |
| CUDA | submission.cu, following templates/template.cu | The queue only accepts submission.py, so run python wrap_cuda_submission.py <SUNet ID> first to wrap the CUDA code |
Every problem folder has templates/ (template.py, template.cu) and test_cases/test.txt. The FlashAttention problem also ships template_triton.py, a Triton kernel skeleton with online softmax.
The README calls Triton and TileLang "modern AI frameworks that provide tile-based abstractions." If you read L13 on separating algorithm from schedule, these tools are that idea applied to GPU kernels. For Triton basics, start with the CS336 kernels and Triton post.
The five problems
Sizes below come from each problem's test_cases/test.txt or README.
| Problem | What it computes | Test size | Difficulty the README points to |
|---|---|---|---|
| Histogram | Multi-channel histogram: for an integer array of shape [length, num_channels], count each bin per channel | length 1,048,576, 512 channels, 256 bins | Many threads competing for the same bins cause atomic contention and poor access patterns |
| 1d-occupancy-decoder | MLP embedder → cross-attention → LayerNorm → output projection, with module and input sizes taken from Roblox's open-source Cube3D | 250,000 queries, 1,024 latents, width 768, 12 heads | Queries vastly outnumber keys/values; everything is float16 except softmax, which must be float32 |
| FlashAttention | softmax(QKᵀ/√D)V on FP16 inputs | Three shapes; the remote benchmark uses only the largest, (4, 64, 8192, 128) | Standard attention is memory-bound on an H100 |
| 3D Heat Equation – RK4 | An 8th-order, 25-point stencil Laplacian inside classical RK4 time stepping | 600³ grid, 10 steps | Correctness requires rtol and atol of 1e-6; the 4-cell boundary stays fixed |
| SwiGLU | Swish(xW + b) ⊙ (xV + c) | batch 256, in_features 2048, hidden 4096 | The README gives only the definition and shapes; you find the bottleneck yourself |
The first question for each problem
These are not solutions. They are the questions to ask after reading each README, tied to concepts from earlier in the course.
Histogram: how many threads write the same bin at once? That is the contention from L6 Locality and communication. The README points toward shared memory and synchronization.
1d-occupancy-decoder: the README includes the TA's own attempts. The PyTorch version takes about 7.4 ms. Porting both the MLP embedder and cross-attention to Triton without fusion made it 19% slower; porting only the embedder made it 2% faster. The TA says the cross-attention port took five to six hours. That table is a lesson by itself: a layer-by-layer rewrite without fusion often loses to a heavily tuned library.
FlashAttention: the README names three key ideas: tiling (load small blocks of Q, K, V into SRAM), fusion (never write the N×N score matrix to HBM), and online softmax. In the README's reference table, PyTorch's built-in FlashAttention runs the largest shape in about 28 ms on an H100, 4.5× faster than the naive reference. That tells you roughly what a library already achieves.
RK4: each time step computes the Laplacian four times, and each Laplacian reads 25 neighbors. The README's baselines are about 1458 ms for PyTorch, 317 ms for naive Triton, and 148 ms for naive CUDA. Decide first whether the intermediate k₁ through k₄ need to go back to global memory. The README itself names custom memory access patterns, shared memory, and kernel fusion.
SwiGLU: two matrix multiplies share the same input x, followed by elementwise work. Profile first to see whether the time goes to the matmuls or the elementwise part, then decide whether to fuse.
How much FlashAttention you need for this problem
The minimum to do the assignment:
- Standard attention computes the N×N score matrix, writes it to HBM, reads it back for softmax, then multiplies by V. At sequence length 8192, reading and writing that matrix is the main cost.
- FlashAttention splits Q into row blocks. Each block sweeps over the blocks of K and V, computing entirely in SRAM.
- Softmax needs each row's maximum and sum. Online softmax updates both as it visits each K block and rescales the output accumulated so far.
For the full derivation, the backward recomputation trick, and how FA2 through FA4 changed with the hardware, read the CMU 11-868 FlashAttention post. The README recommends the original FlashAttention paper and UW CSE 599M's notes From Online Softmax to FlashAttention.
Running code: the job queue and local development
The course workflow has four steps, and the first three are tied to a Stanford identity:
- Install popcorn-cli (prebuilt Linux and macOS binaries are in the repo's
binary/folder, or build it with Rust) - Run the
setup.shposted on Ed to configure the server connection - Register with
popcorn-cli register --sunet-id ... --nickname ...; the team name cannot be changed afterward - Submit with
popcorn-cli submit --leaderboard <problem> --mode <mode> submission.py
There are four --mode values:
| Mode | What it does |
|---|---|
test | Checks correctness only |
benchmark | Measures runtime without posting to the leaderboard |
leaderboard | Measures runtime and posts to that problem's leaderboard |
profile | Runs Nsight Compute and returns a summary; the full .ncu-rep is downloadable from the job page |
The profile summary lists Compute_Throughput, SM_Busy, L1/L2/DRAM throughput, cache hit rates, and traffic between levels. The README shows an RK4 heat_step_kernel example: 7.10 µs, with Compute_Throughput at 18.45% and DRAM_Throughput at 4.45%, then asks why the kernel isn't using the GPU well yet. RK4 profiles report only the first 20 kernels.
Local development needs no SUNet. On any CUDA-capable NVIDIA GPU, add problems/ to PYTHONPATH and run this in a problem folder:
python ../eval.py <test/benchmark/profile> test_cases/test.txt
profile mode needs ncu, the ncu_report Python package, and permission to read GPU performance counters; the README has step-by-step commands. Its example local machine is an AWS g6.xlarge (NVIDIA L4). It also warns that a smaller GPU is fine for correctness and exploration, but you should tune for the H100 and report analysis on the H100, because the best decision for one processor may not be best for another.
What you can do from outside Stanford
- H100 queue and leaderboard: no. Registration needs a SUNet ID, and
setup.shis only posted on the course Ed. So is the leaderboard link. - Problems, templates, references, eval.py: all public. With any NVIDIA GPU you can run correctness tests and benchmarks.
- H100 numbers: rent one yourself. Without an H100, your numbers aren't comparable to the README's reference values, so name your GPU in your log.
For outside readers, then, this assignment is A3 minus the grading environment: the materials are complete, but the course H100s and classmates' leaderboard are missing. The access levels are defined in the global AI/CS course map.
Grading: the work log
| Score | Criteria |
|---|---|
| 80 | Minimal effort, but the log shows course concepts applied over a few optimization steps, with some speedup |
| 95 | Uses course concepts to interpret runtime and profiler results, argues for what to try next, and explores a reasonable set of options; a decent final number also counts as evidence |
| 95–110 | Everything for 95, plus an impressive result; reaching 100 may mean a high leaderboard position, and points above 100 are case by case |
The README reserves the right to go below 80 if the minimal-effort bar isn't met.
The log has three parts:
- The steps you took. For each step: how the code is structured (submit the code, but also describe it at a high level, such as "we blocked the outermost loop and mapped blocks to CUDA thread blocks"), its runtime, which statistics you looked at, what you concluded, what you think limits performance, and what that suggests changing next. Skip the small tweaks. Aim for the level at which you discussed PA2 through PA4 with CAs in office hours.
- Why you stopped. Running out of time is an acceptable answer. If you stopped because the profile showed little left to gain, explain how you decided.
- Whether LLMs helped. Did you use them to write code, brainstorm, or read profiles, and was it useful? The README says iterating with an LLM or running a sequence of prompt-engineering steps is also a good way to do the assignment, as long as the log records your reasoning and prompts. If an LLM hands you a result on the first prompt that you can't improve, the README asks you to contact the staff; one option is to attempt a second problem.
The hand-in is a single .zip with handin.pdf and the code for your key steps.
The README describes the loop in four steps: run, measure, form a hypothesis from your understanding of the code plus the data (increase parallelism, reduce memory traffic, hide memory latency, alleviate contention), and change the code to test it. Those four directions read almost like the table of contents for the first half of CS149: parallelism and latency hiding in L2, bandwidth in L3, contention in L6, and the CUDA memory hierarchy in L7.
Try this: pick a problem and don't write a kernel yet. Run the baseline with eval.py profile, paste the summary into your notes, and finish this sentence: "I think it is limited by ___, because of this number: ___." That sentence is step one of your work log.
What this post can and cannot confirm
Confirmed: the asst5-kernels README, the five problem READMEs and test_cases/test.txt files, the template file list, and the assignment title and due date on the course home page. Performance numbers in the READMEs were measured by staff on their own hardware (the FlashAttention table's small and medium shapes on an RTX 5090, the large shape on an H100). This post quotes them without rerunning anything.
Not confirmed: the contents of setup.sh on Ed, actual leaderboard results, queue wait times on the H100s, and the finer points of how staff grade work logs.
One more thing to watch: the SwiGLU README calls arXiv 1710.05941 "the original SwiGLU paper," but that is Ramachandran, Zoph, and Le's 2017 "Searching for Activation Functions," which introduced the Swish activation. The paper that put Swish inside a GLU gate and named it SwiGLU is Noam Shazeer's 2020 "GLU Variants Improve Transformer" (arXiv 2002.05202). The problem's formula, Swish(xW + b) ⊙ (xV + c), draws on both: look up Swish in the first and the SwiGLU structure in the second.
Further reading: the full FlashAttention story in CMU 11-868 L21; Triton's programming model in CS336 kernels and Triton; another take on GPU programming in CMU 11-868 GPU programming.
Series: previous L13 DSLs and AI-driven optimization | next L14 Cache coherence: MSI, MESI, and false sharing | Series overview
References
- Stanford CS149 Fall 2025 home page and assignment list
- Assignment 5 README (stanford-cs149/asst5-kernels)
- Histogram problem
- 1d-occupancy-decoder problem
- FlashAttention problem
- 3D Heat Equation – RK4 problem
- SwiGLU problem
- Searching for Activation Functions (arXiv 1710.05941, source of Swish)
- GLU Variants Improve Transformer (arXiv 2002.05202, source of SwiGLU)
- GPU MODE popcorn-cli and kernelbot (the grading infrastructure PA5 uses)
- Triton documentation
- TileLang documentation
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (arXiv 2205.14135)
- From Online Softmax to FlashAttention (UW CSE 599M notes, PDF)
Loading...