Skip to content

CS149 L10: Why General-Purpose Processors Waste Energy — Hardware Specialization, Tensor Cores, TPU Systolic Arrays, and Dataflow Architectures

Sep 30, 20261 min
TL;DRL10 starts from one equation: when power is capped, performance can only improve by spending fewer joules per operation, and a general-purpose processor spends most of its energy fetching, decoding, and moving data rather than computing. The slides' rule of thumb is that GPUs give about 10x better perf/watt than CPUs and fixed-function ASICs can reach 100–1000x. The lecture then judges the H100's Tensor Cores and TMA, Google's TPU systolic array, and reconfigurable dataflow architectures against the same checklist: tiled tensors, asynchronous compute and memory, and compute units talking directly to each other.

🌏 中文版

This guide follows the Fall 2025 edition of CS149. It is post 13 in the Reading Stanford CS149 series and covers Lecture 10 from October 23, Hardware Specialization. The official slide PDF has 71 slides.

Fall 2025 recordings live only on Stanford Canvas. The closest public video is 2023 Lecture 18: Hardware Specialization, but it only supplements the first half. Compared with the 2023 course site's slides on the same topic, the opening material (energy constraints, H.264, FFT, DSPs, Anton, FPGAs, efficiency rules of thumb) is all in the 2023 deck. The 2023 second half covered the Spatial accelerator-design language, streaming execution, and how DRAM works. The 2025 deck replaces that with GPU Tensor Cores, the TPU systolic array, and dataflow architectures. This guide follows the 2025 slides. The course overall is A3 (enough to self-study); gaps are listed in the series overview.

The previous post ended on a question: GPUs run DNNs well, but are they the ideal platform? This lecture answers it. This post also sets up the accelerator vocabulary (Tensor Core, systolic array, TMA, dataflow architecture) that the next post uses without re-explaining.

Why specialize: energy

Slides 2–4 reframe the problem as energy. Phones are limited by battery life and fanless heat dissipation. Supercomputers and data centers are limited by power and cooling because of their sheer scale (hundreds of thousands of CPUs and GPUs). And AI demand is growing exponentially.

Slide 5's equation is the backbone of the lecture:

Power = (Ops / second) × (Joules / Op)

Power is a fixed cap, so going faster means spending less energy per operation. The slide's conclusion: better energy efficiency ⇒ specialization (fixed function). Then it asks how big the improvement from specialization can be.

Where a general-purpose processor spends its energy

Slide 7 anticipates the objection: this whole class has been about using multi-core CPUs and GPUs efficiently, and now they're "inefficient"?

Slide 8 lists what a modern processor does to execute one instruction: read it (address translation, icache access), decode it (translate to micro-ops), check dependencies and pipeline hazards, find a free execution resource, read operands from the register file, move data to the execution unit, do the arithmetic, move the result back, write the register file. Only one step is the actual computation. A review question follows: how does SIMD reduce this overhead for some computations, and what properties must those computations have?

Two measured examples:

  • H.264 video encoding (slide 9, Hameed et al., ISCA 2010): even with SIMD, functional units consume only a small fraction of the energy. The rest of the chart goes to instruction fetch, register access, pipeline control, and the data cache.
  • FFT (slide 10, Chung et al., MICRO 2010): an ASIC matches one CPU core's performance with about 1/1000 of the chip area and about 1/100 of the power. GPU cores are about 5–7x more area-efficient than CPU cores.

The specialization spectrum

Slides 11–17 walk through options more specialized than a CPU:

  • DSPs (slide 11): still programmable, but with simpler instruction-stream control. They use complex instructions (SIMD, VLIW) to do many operations per instruction and amortize control. The example is Qualcomm's Hexagon DSP, used for modem, audio, and imaging on Snapdragon chips; the innermost FFT loop runs 29 "RISC" ops per cycle. VLIW (very long instruction word) means one instruction specifies several different operations, in contrast to SIMD's same operation on many data.
  • Anton (slide 12): D. E. Shaw Research's molecular dynamics supercomputer. Anton 1 (2008) has 512 ASICs for particle-particle interactions, a throughput-oriented FFT subsystem, and a low-latency network built for N-body communication patterns. The slide says Anton 3 (2025) is about 20x faster than a contemporary GPU.
  • TPU (slide 13): Google's deep learning processor, after which top architecture conferences filled with DNN accelerator papers.
  • FPGAs (slides 14–17): a middle ground between an ASIC and a processor. The chip provides an array of logic blocks and interconnect, and programmer-defined logic is implemented directly on it. The basic unit is a programmable lookup table (LUT); a 6-input LUT in a Xilinx Virtex-7 is a 64-entry table, and chaining eight LUT6s gives a 40-input AND. Modern FPGAs devote a lot of area to hard blocks (SRAM, multiplier DSP blocks, even ARM or RISC-V CPUs) and are programmed in hardware description languages like Verilog. AWS offers cloud FPGAs through EC2 F1/F2.

Slide 18's rules of thumb, compared with high-quality C code on a CPU:

Platformperf/watt improvementAssumption
Throughput-oriented processors (GPU cores)About 10xCode maps well to wide data-parallel execution and is compute bound
Fixed-function ASICCan approach 100–1000x or moreCompute bound and not floating-point math

Slide 19 (design credited to Pat Hanrahan) lays these out as a spectrum: CPU, GPU, programmable DSP, domain-specific accelerator, FPGA, ASIC. Efficiency rises and programmability falls from left to right. CPUs are easiest to program. Domain-specific accelerators are programmable only within a limited domain, through DSLs (for example, DNNs). FPGAs are hard to program, and making them easier is an active research area. ASICs aren't programmable and cost tens to hundreds of millions of dollars to design, verify, and build.

What an ideal AI accelerator looks like

Slides 21–22 return to last lecture's question. GPUs offer high FLOPS and cuDNN, but if the main operation in AI is matrix multiply, is a general-purpose processor needed?

Slide 23 lists traits of an ideal AI accelerator: high peak TFLOPS and energy efficiency, high memory bandwidth, easy to program for high performance, and able to hit the performance bound on both compute-bound and bandwidth-bound models.

Slides 24–28 build up an "ideal features" table. Two ideas set it up: asynchronous execution (later loads, ops, and stores start before earlier ones finish), and the fact that AI models are dataflow graphs (slide 25: GEMM → Pool → GEMM → SoftMax → Sum). Slide 26 names the crux: GEMM computation is cheap; data movement is expensive, in silicon area, watts, and nanoseconds.

FeatureWhy
Tiled tensors (e.g., 16×16, 32×32)Max TFLOPS on GEMM, low instruction overhead
Asynchronous computeOverlap compute and memory access
Asynchronous memory accessOverlap compute and memory access
Asynchronous chip-to-chip communicationOverlap compute, memory, and communication
Compute unit to compute unit communicationFusion and pipelining, streaming dataflow

The last row is the hardware version of the previous post's fusion: intermediates never leave the chip.

Two key numbers: moving data is expensive, and so is control

Slide 31's "ballpark" numbers (sources: Bill Dally and Tom Olson) count only the logical operation, not decode or register reads:

  • Integer op: about 1 pJ; floating-point op: about 20 pJ
  • Reading 64 bits from a small local SRAM 1 mm away on chip: about 26 pJ
  • Reading 64 bits from low-power mobile DRAM (LPDDR): about 1200 pJ

Hence the design rule: always look for ways to reduce data movement.

Slide 32 quantifies the overhead of programmability (instruction stream and control) relative to the operation itself:

  • Half-precision FMA (fused multiply-add): 2000%
  • Half-precision DP4 (vec4 dot product): 500%
  • Half-precision 4×4 MMA (matrix multiply-accumulate): 27%

The principle: amortize instruction-processing cost across the many operations of a single complex instruction. Tensor Cores are the result.

Slide 33 covers number formats (credited to Bill Dally). BF16 has 1 sign bit, 8 exponent bits, and 7 mantissa bits: the same range as FP32 with lower accuracy. FP8 comes as E4M3 (range 0–448) and E5M2 (range 0–57344).

How GPUs moved toward specialization

A100 and H100 Tensor Cores

Slide 35's A100 SM has 64 fp32 ALUs, 32 int32 ALUs, and 4 Tensor Cores. A Tensor Core executes an 8×4 × 4×8 matrix multiply-add A×B + D, with A and B stored as fp16 and accumulation in fp32. The full GA100 has 108 SMs; at 1.4 GHz that's 19.5 TFLOPS of fp32 plus 312 TFLOPS of mixed fp16/32 in the Tensor Cores.

Slides 36–41 cover the H100 (2022): fourth-generation Tensor Cores, TMA, CUDA clusters, up to 80 GB of HBM3, TSMC 4nm, 80 billion transistors. Slide 38 lines up three hierarchies:

CUDA hierarchyCompute hierarchyMemory
GridGPU80 GB HBM / 50 MB L2
ClusterCPC256 KB shared memory per SM
Thread BlockSM256 KB shared memory
ThreadSIMD lane1 KB registers per thread, 64 KB per SM partition

A thread block cluster holds up to 16 thread blocks, each guaranteed to run on a separate SM at the same time. Slide 41's full H100 has 144 SMs. The Tensor Cores (labeled "systolic array MMA" on the slide) deliver 989 TFLOPS fp16; the SIMD units deliver 134 TFLOPS fp16 and 67 TFLOPS fp32. Slide 43's title says it plainly: all the TFLOPS are in the Tensor Cores.

TMA: a unit just for moving data

Slide 40's TMA (Tensor Memory Accelerator) provides special-purpose instructions for data movement. One thread issues a copy descriptor describing a tensor region; hardware generates the addresses, moves the region asynchronously from global to shared memory, and signals a barrier when the copy completes.

B100: "Not your father's CUDA"

Slide 44 lists the specialized features each generation added from V100 to A100, H100, and B100 (FP8, FP4, Transformer Engine, asynchronous copy, Tensor Core sparsity, a decompression engine, and more), and asks what that means for programmers.

Slide 45 gives part of the answer. B100 Tensor Cores are limited by register bandwidth, so tensor data lives in SMEM and TMEM, and single threads issue MMAs: "No more warps!" Programming Tensor Cores becomes: allocate TMEM and descriptors with tcgen05.alloc; prefetch and stream tiles with TMA (cp.async.bulk.tensor, coordinated with mbarrier); launch async MMAs with tcgen05.mma and tcgen05.commit; order and retire with tcgen05.fence. The slide's tagline: "Not your father's CUDA." Slide 46 then lists DSLs for GPU AI kernels, such as Cute-DSL (CUTLASS in Python) and Mosaic GPU.

How close GPUs are to ideal

Slide 47 puts GPUs back into the ideal-features table: tiled tensors ✅; asynchronous compute ✅ (mma_async); asynchronous memory access ✅ (TMA + TMEM); compute-unit-to-compute-unit communication ❓, only partly covered by thread block clusters.

Google's TPU and the systolic array

Slide 49 shows a lineup of AI accelerators: AWS Trainium 2, Google TPU3, Apple Neural Engine, an Intel inference accelerator, SambaNova Cardinal SN10, the Cerebras Wafer Scale Engine, and an Ampere GPU with Tensor Cores.

Slides 50–51 look at TPU v1 (figures from Jouppi et al. 2017). Arithmetic units take about 30% of the chip, and control takes very little area. There are only five key instructions: read host memory, write host memory, read weights, matrix_multiply/convolve, and activate.

Slides 52–58 animate the systolic array. Take y = Wx: each PE in a 4×4 grid holds one weight. Elements of x enter from the left on a diagonal, one cell per clock. Each PE does one multiply and passes its partial sum down to its neighbor, and results land in 32-bit accumulators at the bottom. For a matrix multiply Y = WX, columns of X flow through wave after wave, and the slide notes you need multiple accumulators to hold the output columns.

Slide 59's comparison:

SIMDSystolic array
DataflowControl-driven (instructions)Data-driven (wavefront)
Data reuseLimitedTemporal and spatial
CommunicationGlobal (register/memory)Local (neighbor PEs)
ControlCentralizedDistributed
Efficiency (perf/mm², perf/W)MediumVery high
Slides 60–63: computing a big matrix on a small array

The example is A = 8×8, B = 8×4096, C = 8×4096, assuming 4096 accumulators. Over four animated slides, A is split into small blocks loaded into the array in turn, columns of B stream through, and partial sums accumulate across the 4096 accumulators. The point: the array size is fixed, so large matrices are handled by blocking and rotating through accumulators. It's the blocked GEMM from the previous post, scheduled in hardware.

Slides 64–65 show the TPU's perf/watt comparison and how TPUs evolved across generations. Slide 66 cites Sara Hooker's Hardware Lottery: a research idea may win because it suits the dominant available software and hardware, not because it's universally better. The slide draws a loop: the TPU is good at dense matrix multiply (labeled "OI ∝ n": the bigger the matrix, the more work per unit of data moved), transformer models fit that well, so hardware specializes even further for matrix multiply.

Dataflow architectures

Slides 67–69 return to "AI models are dataflow graphs." If so, why not make the hardware dataflow too? The example is Plasticine (Prabhakar, Zhang, et al., ISCA 2017): a chip tiled with PCUs (pattern compute units), PMUs (pattern memory units), and switches, which lays out GEMM plus parallel patterns like map, filter, and reduce directly on the chip.

Slide 69 puts it back into the ideal-features table and adds two advantages GPUs lack: no instructions ⇒ no instruction fetch/decode overhead, and extreme asynchrony: no sequential instruction execution. Slide 70 shows FlashAttention on a dataflow architecture. QKᵀ, Mask, Softmax, Dropout, and ×V each get a group of PCUs and PMUs, and tiles flow through them like an assembly line, called a metapipeline. How to program this kind of hardware is the subject of the next post.

Wrap-up: three things specialized hardware shares

Slide 71's summary. Specialized hardware for DNNs:

  1. has many arithmetic units;
  2. has customized or configurable datapaths that move intermediate values directly between processing units, which means scheduling the computation by laying it out spatially on the chip, at multiple granularities;
  3. has large amounts of on-chip storage for fast access to intermediates.

Vocabulary for the next few posts

TermIn one line
ASICFixed-function circuit: most efficient, not programmable
FPGAReconfigurable logic, between an ASIC and a processor
Tensor CoreA GPU unit that does a small matrix multiply-add in one instruction
TMAA GPU unit dedicated to moving tensor blocks asynchronously
Systolic arrayA multiply-add grid where data flows between neighboring PEs; the heart of the TPU
Dataflow architectureLays the computation graph on chip; units pass data directly, no instruction stream

Something to try tonight: use slide 31's numbers to estimate a matrix multiply of two 4096×4096 fp16 matrices. If every input is read from DRAM once, how much energy goes to data movement versus arithmetic? Then think about which side blocking saves.

Further reading: for how TPUs are used with JAX and Pallas in an LLM systems course, read CMU 11-868 TPU, JAX, and Pallas. For GPUs and TPUs from the LLM training side, read CS336 GPUs and TPUs.

Series navigation: previous L9 Running DNNs efficiently on GPUs | next L11 Programming systems for specialized hardware | Series overview

References