Skip to content
All tags

#model-compression

3 posts

MIT 6.5940 Lecture 4: Per-Layer Pruning Ratios, Fine-Tuning, and Hardware Support for Sparsity

MIT 6.5940 Lecture 4 finishes the pruning unit. Per-layer ratios come from sensitivity analysis, AMC (reinforcement learning), or NetAdapt (step-by-step with a lookup table). Fine-tuning uses 1/10 to 1/100 of the original learning rate, and iterative pruning pushes AlexNet from 5x to 9x. EIE, NVIDIA 2:4 sparsity, and TorchSparse/PointAcc show that sparsity only turns into speed with system support.

MIT 6.5940 Lecture 5: Number Formats, K-Means Quantization, and Linear Quantization

MIT 6.5940 Lecture 5 starts from one fact: an 8-bit integer add uses 30x less energy than a 32-bit float add. It reviews the bit layouts of INT, fixed point, FP32/FP16/BF16, FP8, and FP4, then covers two quantization methods. K-means quantization saves storage only, since computation stays in floating point. Linear quantization, r = S(q − Z), turns matrix multiplication, fully connected layers, and convolutions into integer arithmetic.

MIT 6.5940 Lecture 6: Quantization II — PTQ Granularity and Clipping, QAT and STE, Binarization, Mixed Precision

Lecture 6 is about what to do when quantization costs you accuracy. First, without retraining: use finer scale granularity (per-channel, group, MX), clip outliers (EMA, calibration batches, MSE, KL), and round smarter (AdaRound). If that isn't enough, retrain: QAT keeps a full-precision copy of the weights, runs fake quantization in the forward pass, and uses the STE to pass gradients straight through. In the whitepaper table the slides cite, MobileNetV1 drops to 0.1% accuracy under per-tensor INT8 PTQ and recovers to 70.7% with per-channel QAT, against a 70.9% float baseline. The last two sections cover 1–2 bit binary and ternary networks, and HAQ, which uses reinforcement learning to assign a bit width to each layer.