Skip to content

MIT 6.5940 Lecture 4: Per-Layer Pruning Ratios, Fine-Tuning, and Hardware Support for Sparsity

Sep 30, 20261 min
TL;DRMIT 6.5940 Lecture 4 finishes the pruning unit. Per-layer ratios come from sensitivity analysis, AMC (reinforcement learning), or NetAdapt (step-by-step with a lookup table). Fine-tuning uses 1/10 to 1/100 of the original learning rate, and iterative pruning pushes AlexNet from 5x to 9x. EIE, NVIDIA 2:4 sparsity, and TorchSparse/PointAcc show that sparsity only turns into speed with system support.

🌏 中文版

Version note: This post is based on the MIT 6.5940 Fall 2024 course page, the most recent complete offering. Fall 2025 was not offered because Song Han was on sabbatical; the series overview explains the choice. The main source is the Lecture 4 slides, Lec04-Pruning-II.pdf (119 pages; page numbers below are PDF pages). The recording is linked too, but every claim here rests on the slides. Facts were checked against the official materials on 2026-09-30. Access level: Fall 2024 is A3 (slides, recordings, and labs all public); Fall 2026 is A2 (in progress).

Series: previous Lecture 3: where to prune, at what granularity, by what criterion | next Lab 1: fine-grained vs. channel pruning | Series overview

Lecture 3 answered two questions: what shape to prune, and which weights to pick. Pruning a real model raises three more. How much should each layer lose? How do you recover the accuracy you lost? And will the hardware actually run the sparse matrix any faster?

Lecture 4 covers those three. Page 2 splits the pruning unit into five questions. Lecture 3 took the first three; this lecture takes "Determine the Pruning Ratio" and "Fine-tune/Train Pruned Neural Network", then adds a long section on system and hardware support.

Pruning as an optimization problem

Page 4 writes pruning down formally:

$$ \arg\min_{W_P} L(x; W_P) \quad \text{s.t.} \quad |W_P|_0 < N $$

$L$ is the training objective, $W_P$ the pruned weights, $|W_P|_0$ the count of nonzeros, and $N$ the budget of nonzeros you allow. The constraint caps the total. It says nothing about how that total splits across layers, and the first section of the lecture fills that gap.

How much to prune in each layer

Non-uniform beats uniform

Pages 10–11 restate a result: shrinking every layer's channels by the same factor (uniform shrink) loses to pruning each layer by a different ratio. The slide cites the accuracy-vs-latency curve from AMC (He et al., ECCV 2018), where the pruned models reach higher accuracy at the same latency.

So how do you set the ratios?

Method 1: sensitivity analysis

Page 12 gives the intuition. Layers differ in how much pruning they tolerate. Some are sensitive (the first layer, for example) and some are redundant. Page 16 lays out the procedure, using VGG-11 on CIFAR-10:

  1. Pick a layer $L_i$.
  2. Prune it at each $r \in {0, 0.1, 0.2, \dots, 0.9}$.
  3. Record the accuracy drop $\Delta Acc_r$ at each ratio.
  4. Repeat for every layer.

With one curve per layer, draw an accuracy threshold $T$ and give each layer the largest ratio that stays above it.

Pages 25–26 ask whether this is optimal. The slide answers "Maybe not": it touches one layer at a time and ignores how layers interact. The next method goes after that gap.

Method 2: AMC, hand the ratios to reinforcement learning

Page 28 states the goal as a "push-the-button" solution that doesn't need an expert in both ML and hardware. AMC treats layer-by-layer ratio selection as a reinforcement learning problem. Page 31 lists the setup:

ComponentSetting
StateLayer index, channel count, kernel size, FLOPs, and other features
ActionA continuous value $a \in [0, 1)$: the layer's pruning ratio
AgentDDPG, because it supports continuous actions
Reward$-\text{Error}$ if the constraints are met, $-\infty$ otherwise

Latency constraints use a pre-built lookup table instead of on-device measurement every time.

The table on page 34 is the concrete evidence. The baseline MobileNet has 569M MACs, 70.6% top-1, and 119.0 ms. AMC at 50% FLOPs gets 285M MACs, 70.5%, and 64.4 ms. The hand-designed comparison, MobileNet at 0.75 width, lands at 325M MACs, 68.4%, and 69.5 ms. Latency was measured with TF-Lite on a Samsung Galaxy S7 Edge, single core, batch size 1.

Method 3: NetAdapt, trim a little and keep the best layer

NetAdapt (Yang et al., ECCV 2018) takes a rule-based, iterative path (page 35). It also searches for per-layer ratios that meet one global resource budget, such as latency or energy. The loop on page 41:

  1. Each round sets a latency reduction target $\Delta R$ (chosen by hand).
  2. For every layer, prune just enough to save $\Delta R$ (estimated from a lookup table), fine-tune briefly for 10k iterations, and measure accuracy.
  3. Keep the layer whose pruned version scored highest, and prune it for real.
  4. Repeat until the total latency meets the budget, then fine-tune for a long run to recover accuracy.

Page 42 points out a side effect: every round yields a model, so you end up with a whole family of models at different costs, one per iteration.

Recovering accuracy after pruning

Page 45: the higher the pruning ratio, the more accuracy drops. Fine-tuning the pruned network recovers accuracy and lets you push the ratio further. The slide gives a working number: the fine-tuning learning rate is usually 1/100 to 1/10 of the original.

Pages 46–53 cover iterative pruning. One round is "prune, then fine-tune", and each round raises the target sparsity a bit. Page 53 cites Han et al. (NeurIPS 2015): on AlexNet, iterative pruning raised the pruning ratio from 5x to 9x compared with one aggressive step.

Page 54 adds regularization, a term in the loss that penalizes nonzero parameters and pushes parameters toward smaller values:

  • L1: $L' = L(x; W) + \lambda |W|$
  • L2: $L' = L(x; W) + \lambda |W|^2$

The slide's examples: magnitude-based fine-grained pruning applies L2 to the weights, and Network Slimming (Liu et al., ICCV 2017) applies smooth-L1 to the channel scaling factors.

Sparsity needs system support to get faster

This is the longest section of the lecture (pages 56–117). Page 57 lists three case studies, one for each kind of sparsity:

CaseSparsity it exploits
EIEWeight sparsity plus activation sparsity
NVIDIA Tensor CoreM:N weight sparsity
TorchSparse and PointAccActivation sparsity (sparse convolution on point clouds)

EIE: the first accelerator for sparse, compressed models

Page 60 calls EIE (Han et al., ISCA 2016) "The First DNN Accelerator for Sparse, Compressed Model". It exploits three things at once:

  • Sparse weights: 90% static sparsity, for 10x less computation and 5x less memory.
  • Sparse activations: 70% dynamic sparsity, for another 3x less computation.
  • Weight sharing: 4-bit weights, for 8x less memory.

Pages 61–75 walk through how EIE partitions the sparse matrix across processing elements (PEs), its dataflow, and each PE's microarchitecture. These pages read best alongside the recording.

Page 80 is worth copying down, because the slide lists the pros and cons itself:

  • Pros: special-purpose hardware can make sparse operations cost-effective for matrices up to 50% dense. EIE skips both zero weights and zero activations. It supports fine-grained sparsity, which allows higher pruning ratios. It stores 4-bit weights and decodes them to 16-bit for 16-bit arithmetic, and the slide notes that this W4A16 approach is reborn for LLMs in GPTQ, AWQ, llama.cpp, and MLC LLM.
  • Cons: it doesn't map easily onto arrays of vector processors (the fix is structured N:M sparsity). Control flow and storage add overhead (the fix is coarse-grained sparsity). It supports only FC layers. It fits everything in SRAM, which is practical for TinyML but not for LLMs.

Page 81 sums up the section in one principle: the first principle of efficient AI computing is to be lazy. Avoid redundant computation, quickly reject the work, or delay the work.

M:N sparsity: fine-grained sparsity that GPUs can use

NVIDIA's answer to EIE's first con is M:N sparsity. Pages 83–85 cite the 2:4 format from Mishra et al. (arXiv 2021): keep two nonzeros out of every four consecutive weights. Storage packs the nonzeros to the left, so an $R \times C$ matrix becomes $R \times C/2$ values plus 2-bit index metadata.

Page 86 shows how this maps onto a Tensor Core. The sparse matrix A shrinks from $M \times K$ to $M \times K/2$. The hardware uses the indices to pick matching elements from the dense matrix B, so only two of every four multiplications happen.

Does accuracy suffer? The ImageNet top-1 table on page 87 says barely. ResNet-50 goes from 76.1 dense FP16 to 76.2 with 2:4 sparse FP16, for example.

TorchSparse and PointAcc: when the input is already sparse

The third case has nothing to do with pruning: inputs like point clouds are sparse to begin with. Page 90 contrasts the two convolutions. A conventional convolution dilates the nonzeros outward; a sparse convolution does not.

TorchSparse (Tang et al., MLSys 2022) splits sparse convolution into gather, matrix multiply, and scatter-accumulate (page 105). The core trade-off is on pages 106–109:

  • One matmul per kernel offset wastes no computation, but launches many kernels and leaves the device underused.
  • Treating it as a dense convolution is the most regular, but computes far more.
  • Grouping offsets into batched matmuls spends a little extra computation for regularity, and the grouping strategy can be searched per model and dataset (adaptive grouping).

On the hardware side, PointAcc (Lin et al., MICRO 2021) uses merge sort in a mapping unit to find input-output pairs for sparse convolution (pages 115–116). Page 117 charts its speedup and energy savings against platforms including the RTX 2080Ti and TPU v3.

Fall 2026 comparison

Lecture 4 on the Fall 2026 course page (September 22) already has slides and a recording. I compared the text layers of the two PDFs. Both have 119 pages. The only differences are the cover image, the footer's new course name ("TinyML and Efficient AI Computing"), and whitespace. This post applies to both offerings.

The lab is what changed. Fall 2024 released Lab 1 (Pruning) with Lecture 4. Fall 2026 released a different Lab 1, GPU Basics, on the same day. If you want pruning practice, Fall 2024's Lab 1 is the only option, and the next post breaks it down.

What to do after reading

  1. Answer the two questions in the summary on page 118. How do you find pruning ratios automatically? What system support does each granularity need? If you can't, go back to pages 26 and 57.
  2. Map "non-uniform ratios", "fine-tuning", and "hardware support" onto a model you actually run. Does your inference hardware support 2:4 sparsity? If not, fine-grained pruning saves storage there but not time.
  3. One thing you can do tonight: open Fall 2024 Lab 1, run it up to the sensitivity scan cell, and see how different the per-layer curves are.

Further reading

References