Skip to content
All tags

#pruning

3 posts

MIT 6.5940 Lab 1: Nine Questions on Fine-Grained and Channel Pruning

MIT 6.5940 Fall 2024 Lab 1 is one Colab notebook with 9 questions worth 100 points. Questions 1–5 apply magnitude-based fine-grained pruning and a sensitivity scan to a VGG on CIFAR-10, and require a model at 25% of its original size with over 92.5% accuracy after fine-tuning. Questions 6–8 cover channel pruning, Frobenius-norm channel ranking, and measured speedup; Question 9 compares the two. Fall 2026 has no pruning lab.

MIT 6.5940 L3 Pruning I: Where to Prune, How Fine, and by What Criterion

Pruning removes unimportant weights or neurons from a neural network. The goal is written as minimizing loss subject to at most N nonzero weights. Lecture 3 of 6.5940 handles two of the decisions involved. First, granularity: from fine-grained pruning, which can remove any element, to channel pruning, which removes whole channels. The more regular the pattern, the easier it is to speed up on existing hardware, and the less you can remove. In between, 2:4 sparsity gives up to 2× speedup on NVIDIA Ampere GPUs. Second, criteria: look at weight magnitude, Batch Norm scaling factors, second derivatives, the fraction of zero activations, or how well a layer's output can be reconstructed after pruning.

MIT 6.5940 Lecture 4: Per-Layer Pruning Ratios, Fine-Tuning, and Hardware Support for Sparsity

MIT 6.5940 Lecture 4 finishes the pruning unit. Per-layer ratios come from sensitivity analysis, AMC (reinforcement learning), or NetAdapt (step-by-step with a lookup table). Fine-tuning uses 1/10 to 1/100 of the original learning rate, and iterative pruning pushes AlexNet from 5x to 9x. EIE, NVIDIA 2:4 sparsity, and TorchSparse/PointAcc show that sparsity only turns into speed with system support.