Skip to content

MIT 6.5940 Lab 1: Nine Questions on Fine-Grained and Channel Pruning

Sep 30, 20261 min
TL;DRMIT 6.5940 Fall 2024 Lab 1 is one Colab notebook with 9 questions worth 100 points. Questions 1–5 apply magnitude-based fine-grained pruning and a sensitivity scan to a VGG on CIFAR-10, and require a model at 25% of its original size with over 92.5% accuracy after fine-tuning. Questions 6–8 cover channel pruning, Frobenius-norm channel ranking, and measured speedup; Question 9 compares the two. Fall 2026 has no pruning lab.

🌏 中文版

Version note: This post is based on the Lab 1 Colab notebook from MIT 6.5940 Fall 2024. I downloaded the raw notebook on 2026-09-30 and checked question numbers, points, and setup cell by cell. Access level A3: the notebook, pretrained weights, and dataset download are all public. What's missing is official solutions and grading feedback (submission goes through MIT Canvas). This post contains no solutions.

Series: previous Lecture 4: per-layer pruning ratios, fine-tuning, and hardware support | next Lecture 5: number formats, K-means, and linear quantization | Series overview

Lecture 3 and Lecture 4 turn pruning into a pipeline: pick a granularity, pick a criterion, set per-layer ratios, fine-tune, and check whether the hardware can use the result. Lab 1 walks the whole pipeline on a small model, and it deliberately compares the two extremes. One prunes as finely as possible (single weights); the other prunes as coarsely as possible (whole channels).

The notebook opens with five goals. The last two matter most: get a basic understanding of the performance gains from pruning (such as speedup), and understand the differences and trade-offs between the two approaches.

When it runs and what you need

According to the Fall 2024 course page, Lab 1 went out on September 17 (Lecture 4) and was due September 26 (Lecture 7), with the two quantization lectures in between. The collaboration policy: you may discuss answers, but each student hands in their own work and names who they collaborated with.

The notebook's setup does the following:

  • pip install torchprofile (for counting MACs).
  • Checks torch.cuda.is_available() and stops if there's no GPU, telling you to switch the Colab runtime to GPU.
  • Downloads a VGG pretrained on CIFAR-10 (the same model as Lab 0) from hanlab18.mit.edu. That URL still returned HTTP 200 when I tested it on 2026-09-30.
  • Downloads CIFAR-10 with batch size 512.

The notebook says the model is about 35 MiB, and uses that as its opening point: classifying 32×32 images into 10 classes already takes a model this big, which is too heavy for phones and embedded devices.

The nine questions

The notebook has 9 questions worth 100 points, in two sections plus a comparison:

QPointsSectionWhat you do
Q110Fine-grainedRead per-layer weight histograms; describe common traits and how they help pruning
Q215Fine-grainedImplement fine_grained_prune: count zeros, use |W| as importance, find the threshold with kthvalue, build the mask
Q35Fine-grainedSet target_sparsity so the test tensor keeps exactly 10 nonzeros
Q415Fine-grainedRead the sensitivity scan curves: sparsity vs. accuracy, whether all layers are equally sensitive, which layer is most sensitive
Q510Fine-grainedChoose a sparsity for each layer from the sensitivity curves and per-layer parameter counts
Q610ChannelImplement get_num_channels_to_keep and channel_prune
Q715ChannelImplement input-channel importance by Frobenius norm and sort channels by it
Q810ChannelExplain why pruning 30% of channels cuts computation by about half, and why latency drops slightly less than computation
Q910ComparisonPros and cons of each method; which one you'd pick to speed up a model on a smartphone, and why

Only Q2, Q3, Q5, Q6, and Q7 involve code. The rest are short answers based on plots or numbers.

First half: what fine-grained pruning teaches

Q1–Q3: from weight distributions to a mask

The notebook plots a weight histogram for each layer, then moves to magnitude-based pruning. Its definition matches Lecture 3: importance is $|W|$. Given a target sparsity $s$, kthvalue finds the $(#W \cdot s)$-th smallest importance as the threshold, and everything above it stays. The hints for Q2 break this into four steps and even name the PyTorch APIs to use.

Q3 is a quick check. You need the definition of sparsity ($#\text{zeros} / #W$) to find the ratio that leaves exactly 10 nonzeros.

Q4–Q5: be AMC for a day

Next the notebook wraps the pruning function in FineGrainedPruner, which stores each layer's mask so it can be reapplied after weight updates to keep the model sparse.

The sensitivity scan is the procedure from page 16 of Lecture 4: prune one layer at a time, sweep sparsity from 0.4 to 0.9 in steps of 0.1, and record accuracy. The notebook says the cell takes about 2 minutes.

Q5 is the question closest to real work. The notebook also plots the parameter count of each layer and asks you to combine both plots into a sparsity_dict, with a clear pass condition: the pruned model must be 25% of the dense model's size, with validation accuracy above 92.5 after fine-tuning. The hints are two sentences: layers with more parameters should get higher sparsity, and sensitive layers should get lower sparsity.

Then comes fine-tuning: 5 epochs of SGD (lr 0.01, momentum 0.9, weight decay 1e-4) with a cosine schedule, about 3 minutes by the notebook's estimate. This is Lecture 4's point in practice: you need fine-tuning to win the accuracy back.

Second half: what channel pruning teaches

Q6–Q7: prune naively, then learn to choose

The notebook states the selling point of channel pruning: removing whole channels speeds up inference on existing hardware such as GPUs. The pruned weights stay dense, and the output channel count becomes $(1 - \text{sparsity})$ times the original, so this section calls it the prune ratio instead.

Q6 starts with the crudest approach on purpose: prune every layer by 30% and keep the first channels. The notebook says the target is a 2x computation reduction and asks you to think about why 30% gets you roughly there. After running it, you'll see accuracy drop sharply.

Q7 sorts first. Importance is the Frobenius norm of the weights for each input channel; sort, then keep the top $k$. The notebook says sorting improves accuracy only slightly, and fine-tuning (again 5 epochs) does the real recovery.

Q8: measure the real speedup

The final cells compare model size, MACs, and latency before and after pruning. Note the latency setup: the notebook moves the models to the CPU and uses one 1×3×32×32 dummy input, with 20 warm-up runs and 100 timed runs. Q8's two questions ask you to explain the numbers with Lecture 4's ideas, not to recite an answer.

Q9: tie the two halves together

Q9.1 asks you to compare the methods on compression ratio, accuracy, latency, and hardware support (whether they need a specialized accelerator). Q9.2 asks which one you'd use on a smartphone. Neither involves code, but a good answer draws on EIE's pros and cons from Lecture 4 (page 80) and the motivation for M:N sparsity.

Limits of self-study

  • No official solutions and no public autograder. The sanity checks in the notebook (such as test_fine_grained_prune and the channel-sorting check) test function behavior only; nobody grades your short answers.
  • Submission goes through MIT Canvas, so outside learners get no grading feedback. Piazza is limited to enrolled students.
  • The 92.5% bar in Q5 is your only objective checkpoint. For short answers, write against specific slide pages, then check whether you actually used the lecture's concepts.

Fall 2026 comparison

The Fall 2026 course page released Lab 1 with the same Lecture 4 (September 22), but it's lab1_gpu_basics.zip, a lab on GPUs and efficiency fundamentals with no pruning. The Fall 2026 lab list is GPU Basics, Quantization, NAS, Quantization, and LLM deployment on laptop. For pruning practice, this Fall 2024 notebook is the only official material.

What to do after reading

  1. Do Q1–Q5 first. Once Q5 passes, stop and put your sparsity_dict next to the sensitivity curves. Make sure you can justify every number.
  2. After Q8, rerun the latency measurement on the GPU instead of the CPU (the measurement cell already moves the models back to CUDA at the end). Compare the two results, then answer Q9.2.
  3. One thing you can do tonight: run only up to the sensitivity scan, screenshot the curves, mark the layer you think is most sensitive, and check it against Lecture 4's page 12 claim that the first layer is usually the sensitive one.

Further reading

References