CS336 Lecture 5: GPUs Win by Moving Data Less, Not by Making Each Thread Fast
Lecture 5 explains GPUs through SMs, warps, and the memory hierarchy, then unifies common optimization under low precision, fusion, recomputation, coalescing, and tiling. FlashAttention combines those principles for attention.