Skip to content

MIT 6.5940 L3 Pruning I: Where to Prune, How Fine, and by What Criterion

Sep 30, 20261 min
TL;DRPruning removes unimportant weights or neurons from a neural network. The goal is written as minimizing loss subject to at most N nonzero weights. Lecture 3 of 6.5940 handles two of the decisions involved. First, granularity: from fine-grained pruning, which can remove any element, to channel pruning, which removes whole channels. The more regular the pattern, the easier it is to speed up on existing hardware, and the less you can remove. In between, 2:4 sparsity gives up to 2× speedup on NVIDIA Ampere GPUs. Second, criteria: look at weight magnitude, Batch Norm scaling factors, second derivatives, the fraction of zero activations, or how well a layer's output can be reconstructed after pruning.

🌏 中文版

This is post 2 of the Reading MIT 6.5940 series, based on the Fall 2024 edition. The previous post covered how to measure model size and compute. This one starts actually making models smaller.

Official materials covered here:

  • Lecture 3 Pruning and Sparsity (Part I): slides (74 pages), video

All page numbers are PDF page numbers.

Why start with pruning

Page 4 lists the four techniques in the course's first part, "Efficient Inference": Pruning, Quantization, Neural Architecture Search, and Knowledge Distillation. Pruning comes first.

The motivation picks up from the previous lecture's energy table (page 6): one 32-bit DRAM read costs 640 pJ, far more than any arithmetic operation. Fewer weights means less data to move, and energy drops with it.

The slides also include an LLM example (page 5). In the open division of MLPerf Inference v4.1, NVIDIA pruned Llama 2 70B: depth went from 80 layers to 32, and the MLP intermediate dimension from 28,762 to 14,336. On a single H200, offline samples per second rose from 4,488 in the closed division to 11,189, about 2.5×, while keeping 99% accuracy.

The problem statement

Page 9 writes pruning as a constrained optimization problem:

argmin_{W_P} L(x; W_P) subject to ‖W_P‖₀ ≤ N

L is the training objective, W the original weights, and W_P the pruned weights. ‖W_P‖₀ counts the nonzero elements of W_P, and N is how many nonzeros you are allowed. In plain terms: keep only N weights and make the loss as small as possible.

The formula raises four decisions, which also form the outline of Lectures 3 and 4 (page 8):

  1. Granularity: what pattern do you prune in?
  2. Criterion: which synapses or neurons do you remove?
  3. Ratio: what target sparsity does each layer get?
  4. Fine-tuning: how do you recover accuracy afterwards?

This post covers the first two. The other two are in the next lecture.

Pruning needs retraining

Pages 16–20 use experimental curves from Han et al. 2015 to show that pruning is not one cut. The x-axis is the fraction of parameters pruned (40% to 100%) and the y-axis is accuracy loss. The slides add one line per page, three in all:

  • Prune only: the more you prune, the faster accuracy falls
  • Prune, then fine-tune: at the same pruning ratio, the loss is clearly smaller
  • Iterative pruning and fine-tuning: prune a little, train a little, prune again. This reaches the highest ratio with the smallest loss

Page 21 lists results on several classic models:

ModelParams beforeParams afterParam reductionMAC reduction
AlexNet61M6.7M9×3×
VGG-16138M10.3M12×5×
GoogleNet7M2.0M3.5×5×
ResNet5026M7.47M3.4×6.3×
SqueezeNet1M0.38M3.2×3.5×

Parameter reduction and MAC reduction differ because pruned weights do not necessarily sit in the compute-heavy layers. AlexNet's parameters are concentrated in its fully-connected layers, and its compute in its conv layers (see the previous post). So parameters drop 9× while MACs drop only 3×.

Page 22 shows NeuralTalk, an LSTM that generates image captions. With 90% of the weights pruned, its captions are nearly the same as the original model's. Page 24 lists industry hardware support: EIE, ESE, SpArch, SpAtten, and 2:4 sparsity on the A100 GPU (the slide says "2X peak performance, 1.5X measured BERT speedup").

Decision 1: granularity

Starting from a 2D matrix

Pages 29–30 use a 2D weight matrix to show the two extremes:

Fine-grained / UnstructuredCoarse-grained / Structured
What can be prunedAny positionOnly whole rows or blocks
FlexibilityHighLow (a subset of fine-grained)
SpeedupHard, because nonzeros are irregularEasy: the result is just a smaller matrix

Conv layers offer four dimensions

A conv weight has shape [cₒ, cᵢ, k_h, k_w], and the four dimensions allow more ways to prune. Page 33 cites the taxonomy of Mao et al., ordered from irregular to regular:

  1. Fine-grained pruning: any element
  2. Pattern-based pruning: fixed patterns
  3. Vector-level pruning: whole vectors
  4. Kernel-level pruning: a whole k_h × k_w kernel
  5. Channel-level pruning: a whole channel

The slides then look at three representatives.

Fine-grained (pages 35–37). It is the most flexible and usually gives the highest compression, since it can find "redundant" weights anywhere; the table above comes from fine-grained pruning. The drawback, on page 37: it can be accelerated on some custom hardware (such as EIE) but is hard to accelerate on GPUs.

Pattern-based: N:M sparsity (pages 39–42). Remove N of every M consecutive elements. The classic case is 2:4, which is 50% sparsity. The pruned matrix can be stored as the nonzero values plus a 2-bit index per value. The slides cite NVIDIA: the Ampere GPU architecture supports 2:4 sparsity for up to about 2× speedup, and it usually maintains accuracy across many tasks. It is a compromise between regularity and flexibility.

Channel pruning (pages 44–46). Reduce the channel count directly. The result is an ordinary network with fewer channels, so any hardware runs it faster. The cost is a smaller compression ratio. Page 45 compares two approaches: a uniform shrink that cuts 30% from every layer, and channel pruning with a different ratio per layer (for example 0.5, 0.3, 0.7, 0.2). Page 46 cites AMC: at the same latency, per-layer ratios give higher ImageNet accuracy than uniform scaling. How to find those per-layer ratios is the topic of the next lecture.

In short, granularity is a trade-off axis. Finer patterns let you prune more but are harder to accelerate. Coarser patterns accelerate easily but remove less. Which one you choose depends on what your hardware supports.

Decision 2: criteria

A criterion answers "what do we prune?" The principle on page 49: the less important the removed parameters, the better the pruned network performs. The hard part is defining "important."

Page 49 opens with a small example: y = ReLU(10x₀ − 8x₁ + 0.1x₂). If you may remove only one weight, which one? The intuitive answer is 0.1, because it affects the output least. That is the idea behind magnitude-based pruning.

Three criteria for pruning weights

Magnitude-based (pages 50–53). Weights with larger absolute values are more important.

  • Element-wise: importance = |W|. Pruning half of [[3, −2], [1, −5]] keeps 3 and −5
  • Row-wise (structured): importance is the L1 norm of a whole row. For the same matrix, row one is |3| + |−2| = 5 and row two is |1| + |−5| = 6, so row one goes
  • You can also use the L2 norm (√13 vs. √26), or in general the Lp norm (page 53 cites Wen et al. 2016)

The fine-grained part of Lab 1 implements this criterion.

Scaling-based (pages 54–56). This criterion is for filters (output channels) and comes from Network Slimming. Each output channel gets a trainable scaling factor that multiplies its output, and channels with small factors are pruned. Page 56 points out that no extra parameters are needed: Batch Norm's γ is already one scaling factor per channel and can be used directly.

Second-order-based (pages 57–62). This is LeCun's 1989 Optimal Brain Damage: estimate directly how much the loss rises when a weight is removed.

Derivation of Optimal Brain Damage

Expand the change in loss from pruning with a Taylor series:

δL = Σᵢ gᵢ δwᵢ + ½ Σᵢ hᵢᵢ δwᵢ² + ½ Σᵢ≠ⱼ hᵢⱼ δwᵢ δwⱼ + O(‖δW‖³)

Here gᵢ is the first derivative and hᵢⱼ the second derivative (an entry of the Hessian). Optimal Brain Damage makes three assumptions:

  1. The objective is nearly quadratic, so terms of third order and above are dropped
  2. Training has converged, so the first-order term is zero
  3. The errors from deleting each parameter are independent, so the cross terms are zero

What remains is δLᵢ ≈ ½ hᵢᵢ wᵢ². Importance is defined as ½ hᵢᵢ wᵢ², and weights with the smallest error are pruned first.

Page 62 names the practical obstacle: the Hessian is hard to compute.

Two criteria for pruning neurons

Page 63 explains that pruning a neuron is coarse-grained weight pruning. In a fully-connected layer, removing a neuron removes one row of the weight matrix; in a conv layer it removes one channel.

Percentage-of-Zero-based (pages 64–66). ReLU produces many zeros. Measure the average fraction of zeros in each channel's output (APoZ). The smaller it is, the more often the neuron fires and the more important it is. Page 66 works an example with batch 2, 3 channels, and 4×4 maps, and gets APoZ values of 11/32, 12/32, and 14/32 for the three channels, so channel 2 is pruned first. The method comes from Network Trimming.

Regression-based (pages 67–72). The earlier criteria look at the overall loss or the weights themselves. This one looks at a single layer: after pruning, can this layer's output be reconstructed to match the original? Write the original output as Z = XWᵀ, which splits into a sum of contributions from each input channel. Introduce a coefficient vector β of length cᵢ, where β_c = 0 means channel c is pruned. The problem becomes:

argmin_{W, β} ‖Z − Σ_c β_c X_c W_cᵀ‖²_F subject to ‖β‖₀ ≤ N_c

The solution alternates: fix W and solve for β to choose channels, then fix β and solve for W to minimize reconstruction error. The source is He et al., ICCV 2017.

The five criteria side by side

CriterionPrunesLooks atCost
MagnitudeWeights (elements or structures)Lp norm of weightsCheapest
ScalingOutput channelsScaling factors (BN γ works)Must train the factors
Second-orderWeights½ hᵢᵢ wᵢ²Hessian is hard to compute
APoZNeurons / channelsFraction of zero outputsMust run data to collect activations
RegressionChannelsSingle-layer reconstruction errorMust solve an optimization problem

Fall 2026 comparison

The Fall 2026 L3 slides run 71 pages, and the video is up. The three-part structure (intro to pruning, granularity, criteria) matches Fall 2024, and the closing summary is word for word the same. The difference is three pages dropped from the opening: "Today's AI is too BIG," "Efficient Deep Learning Techniques are Essential," and the MLPerf Llama 2 70B pruning case. The charts on the first two already appear in the Fall 2026 L2.

In Fall 2026, Lab 1 became GPU Basics, and the course page and slides disagree on Lab 2's topic (details in the series entry point). To practice this lecture's material, the Fall 2024 Lab 1 is currently the only option.

What you can do tonight

Open the Fall 2024 Lab 1, run Setup and the weight-distribution histograms, and do Question 1: what do the per-layer weight distributions have in common, and why does that help pruning? It needs no code, but if you can answer it, you understand why magnitude-based pruning works. The full walkthrough of Lab 1 is in order 4.

Further reading

Series navigation: previous Why efficiency matters and how to measure model size and compute | next Pruning II: per-layer ratios and hardware support | Series entry point

References