Lecture 10 of 11-868 uses Lei Li's own LightSeq and LightSeq2 as the case study and breaks them into four techniques: fuse every small operation outside matrix multiplication into one kernel, rewrite the LayerNorm and Softmax formulas to cut thread synchronizations, store parameters and gradients in FP16 but compute updates in FP32, and reuse memory based on backward-pass dependencies. The slides report 1.4-3.5x training speedups on WMT14 English-German. There is no recording; this guide works from slide page numbers and the two papers.
CMU 11-868's three GPU lectures answer one question: why does a correct CUDA matmul use only 2.48% of an A100's FP32 compute? L02 covers SMs, warps, and the grid/block/thread hierarchy. L03 covers cudaMalloc, cudaMemcpy, and kernel indexing. L04 uses tiling, coalesced access, and bank-conflict avoidance to bring data closer than global memory, which sits about 500 cycles away.
The first 11-868 assignment has you write four CUDA kernels in src/combine.cu (map 15, zip 25, reduce 25, matmul 30 points), wire them into MiniTorch's Python backend, and finish with a 5-point integration test. Shared-memory optimizations for reduce and matmul are marked Optional. The assignment page says plainly that you need a GPU, and grading uses private test cases.
HW3 has you add softmax loss, Dropout, LayerNorm, and Embedding to the MiniTorch you built in HW1 and HW2, assemble a pre-LN GPT-2 decoder, and train it on IWSLT14 German-English translation. Points: tensor functions 20, basic modules 20, decoder LM 40, translation pipeline 20. Full marks require passing the private tests and a BLEU of about 20±2. The assignment page warns that training alone takes at least 10 hours: one epoch is about an hour on a PSC V100, and you need 10. In Spring 2026 it went out Feb 4 and was due Feb 18.
The fourth 11-868 assignment has you follow LightSeq and hand-write CUDA kernels for attention softmax and LayerNorm (forward and backward), bind them into your own MiniTorch, then swap them into your HW3 Transformer and train for one epoch. Points: Softmax 40, LayerNorm 40, integration 20. The assignment page expects individual kernels to be 3.7x to 15.8x faster, but end-to-end training only about 1.1x faster, because of Amdahl's law. You need an NVIDIA GPU, and the repo has already been changed for Fall 2026.
Lecture 8 asks you to switch mental models: stop thinking about what each worker does and write algorithms as operations on sequences, such as map, fold, scan, segmented scan, gather/scatter, sort, and groupBy. These primitives have efficient parallel implementations, and they turn irregular parallelism into regular parallelism and fine-grained synchronization into coarse synchronization. The price is extra passes over the data, so they are bandwidth hungry.
CUDA's grid, thread block, and CUDA thread are programming abstractions; the GPU implements them with SMs, warps, and a hardware block scheduler. The heart of the lecture is keeping two things apart: the system may run thread blocks in any order, but all threads in one block are guaranteed to be live at once. That is why a block can cooperate through shared memory and __syncthreads(), and why the number of blocks an SM can hold is set by registers and shared memory.
PA3 has three parts: port SAXPY to CUDA and time it two ways, implement find_repeats with an exclusive scan, and write a CUDA circle renderer that is both correct and fast (85 points). The hard part of the renderer is that blending semi-transparent circles doesn't commute, so every pixel must be updated in input order, and the starter code's one-thread-per-circle approach gets neither atomicity nor order right. Written 2 has five graded problems (fusion, SIMD utilization, a barrier instead of locks, data-parallel primitives on graphs, locks in a particle simulation) plus 14 practice problems. Outside Stanford you need your own NVIDIA GPU. No solutions here.
The last programming assignment in CS149 Fall 2025 is open-ended. Pick at least one of five kernels (Histogram, a 1D occupancy decoder, FlashAttention, a 3D heat equation with RK4, SwiGLU) and make it faster than its PyTorch baseline on an H100. You can write CUDA, Triton, or TileLang, and you may use LLMs. There is no speed threshold. The grade depends on a work log that shows what you measured at each step, what hypothesis you formed, and why you stopped. The H100 job queue and leaderboard need a SUNet ID; outside Stanford you can only run eval.py on your own NVIDIA GPU.
Once algorithms have shrunk the model, how much more can the system layer squeeze out? Lecture 11 uses a single matrix multiply to show it: loop reordering gives 12x, tiling 19x (on an Intel Xeon 4114), and a CUDA version runs 94x faster end to end on a 2080Ti. The second half covers TinyEngine's inference tricks: im2col; in-place depthwise, which cuts peak memory from 2×C×H×W to (1+C)×H×W; NHWC for pointwise and NCHW for depthwise; and Winograd, with 2.25x fewer multiplications.
A tour of Karpathy's three teaching repos: nanoGPT (2022, ~300 lines each to reproduce GPT-2 124M), llm.c (pure C/CUDA training), and nanochat (2025-10, one speedrun.sh from tokenizer to WebUI). $100 and 4 hours on 8×H100 buys a chatty model; GPT-2-grade capability is now down to about 2 hours and $48.
TensorRT-LLM is NVIDIA's open-source LLM inference library (Apache 2.0). It offline-compiles model weights and compute graphs into optimized TensorRT engines, then serves them with custom CUDA kernels, in-flight batching, and multi-dimensional parallelism. The cost: NVIDIA GPUs only, compilation takes tens of minutes, and switching models or quantization means rebuilding.
Lecture 5 explains GPUs through SMs, warps, and the memory hierarchy, then unifies common optimization under low precision, fusion, recomputation, coalescing, and tiling. FlashAttention combines those principles for attention.
Lecture 6 turns GPU principles into kernels: benchmark scaling across shapes, profile actual calls and time, then implement GeLU, softmax, reductions, and tiled matrix multiplication in Triton. Speed begins with measuring correctly.
llama.cpp is the most widely used local LLM inference engine, implemented in pure C/C++. It supports CPU, Metal, CUDA, Vulkan, and other backends, and uses the GGUF quantization format to run multi-billion-parameter models on consumer hardware.