🌏 中文版
This post is based on MIT 6.5940, Fall 2024. It is part 7 of the Reading MIT 6.5940 series and turns the quantization material from Lecture 5 and Lecture 6 into code.
Series: previous Lecture 6: PTQ, QAT, and mixed precision | next Lecture 7: NAS search spaces and search strategies | Series overview
Official materials: the Lab 2 Colab notebook. On the Fall 2024 course page, Lab 2 was released on September 26 (Lecture 7) and due October 8 (Lecture 10). The question numbers, points, and wording below follow the notebook itself, checked on 2026-09-30.
Access grade A3, with gaps: the notebook, pretrained weights, and dataset all download directly, and the notebook ships a few verification functions. What you can't get is official solutions or grading: submissions go through MIT's Canvas, so outside learners get no graded feedback. This post doesn't give solutions. It only says what each question asks and which part of the lectures it maps to.
Fall 2026 comparison: The Fall 2026 course page lists Lab 2 as "Quantization," scheduled for release on October 1. When I checked again on 2026-10-01 there was no link yet. I'll add the differences here once it's out.
What the lab wants you to be able to do
The notebook's Goals section lists seven items. Condensed, they come to three:
- Implement K-means quantization and recover accuracy with quantization-aware training (QAT).
- Implement linear quantization and integer-only inference.
- Understand the trade-offs between the two in accuracy, latency, and hardware support.
There are 10 questions: 3 on K-means (Q1–Q3), 6 on linear quantization (Q4–Q9), and Q10 comparing the two.
Setup and starting point
- Model and data: a VGG on CIFAR-10, the same one used in Lab 0. The notebook defines a backbone of 8 conv-bn-relu layers plus a 512→10 linear classifier. Pretrained weights download from
hanlab18.mit.edu. - Runtime: the notebook metadata requests a Colab GPU runtime, and the model is loaded with
.cuda(). - Extra packages:
torchprofile(for MAC counts) andfast-pytorch-kmeans(for K-means). - Built-in checks:
test_k_means_quantize(),test_linear_quantize(), andtest_quantized_fc()check your implementation against fixed small tensors. They're the only automatic feedback outside learners get.
The notebook starts by measuring the FP32 model's accuracy and size, and every later result is compared against those numbers.
Part 1: K-means quantization (Q1–Q3, 30 points)
This part maps to Lecture 5's K-means quantization and the K-means QAT on page 47 of Lecture 6. The notebook's definition: n-bit K-means quantization splits the weights into 2ⁿ clusters and builds a codebook with 2ⁿ FP32 centroids and an n-bit integer labels tensor with as many elements as the original weights. At inference, centroids[labels] reconstructs floating-point weights.
| Question | Points | What it asks |
|---|---|---|
| Q1 | 10 | Complete k_means_quantize(): build a codebook with K-means and reconstruct the weights |
| Q2.1 | 5 | 2-bit quantization renders 4 colors; how many for 4-bit? |
| Q2.2 | 5 | Generalize to n bits |
| Q3 | 10 | Complete update_codebook(): update the centroids from the latest weights |
After Q1, the notebook wraps your function in a KMeansQuantizer class, quantizes the whole model at 8, 4, and 2 bits, and prints size and accuracy. Note that the size calculation ignores codebook storage; the notebook says so explicitly.
The notebook then points out the problem itself: the fewer the bits, the more accuracy drops, so QAT is needed. It gives the centroid gradient as the sum of the weight gradients within a cluster, then explains that the lab, for simplicity, just sets each centroid to the mean of the weights in its cluster. Q3 implements that step.
The fine-tuning loop is provided: fine-tune only if accuracy drops by more than 0.5 percentage points, for at most 5 epochs, using SGD (lr 0.01, momentum 0.9) with a cosine schedule, calling quantizer.apply(model, update_centroids=True) after each training step.
Part 2: Linear quantization (Q4–Q9, 65 points)
This part maps to Lecture 5's linear quantization and Lecture 6's per-channel and integer-inference slides (pages 8–9 and 14–20). It starts from r = S(q − Z) and the n-bit signed integer range [−2ⁿ⁻¹, 2ⁿ⁻¹ − 1].
| Question | Points | What it asks |
|---|---|---|
| Q4 | 10 | Complete linear_quantize(): scale, round, add the zero point; the final clamp to the n-bit range is already written |
| Q5.1 | 3 | Pick the correct scale formula from four options |
| Q5.2 | 4 | Pick the correct zero-point formula from four options |
| Q5.3 | 8 | Complete get_quantization_scale_and_zero_point() |
| Q6 | 5 | Complete bias quantization |
| Q7 | 15 | Complete the integer fully connected layer, quantized_linear() |
| Q8 | 10 | Complete the integer convolution layer, quantized_conv2d() |
| Q9.1 | 5 | Preprocessing that maps inputs from (0, 1) to the INT8 range |
| Q9.2 | 5 | Explain why the quantized model has no ReLU layers |
A few design choices worth knowing up front:
- Weights use symmetric quantization. The notebook plots the weight distributions and notes they're roughly symmetric around 0 (except the classifier), so weights get Z = 0 and S comes from the largest absolute weight. This matches the symmetric linear quantization on Lecture 6 page 14.
- Weights are per-channel. Conv weights have shape (output channels, input channels, kh, kw), and the notebook says extensive experiments show a separate S and Z per output channel works better, so per-channel is built into the provided code.
- Activation ranges are calibrated on one batch of training data. Before Q9, forward hooks record each layer's inputs and outputs on a single batch (batch size 512), and the function you wrote in Q5.3 computes S and Z straight from the min and max. This is the simplest version of Lecture 6 page 32's "calibration batches," with no clipping.
- The Q9 pipeline: first fold BatchNorm into the preceding convolution (the notebook verifies accuracy is unchanged), then record activation ranges, then swap
Conv2dandLinearfor quantized versions.MaxPool2dandAvgPool2dare thin wrappers, because PyTorch modules at the time didn't support INT8, so they temporarily compute in FP32.
The Q7 and Q8 hints give the formula skeleton directly, matching Lecture 6 pages 8–9.
The integer-inference formulas from the Q7/Q8 hints
q_output = (Linear[q_input, q_weight] + Q_bias) · (S_input · S_weight / S_output) + Z_output
Q_bias = q_bias − Linear[Z_input, q_weight]
For convolution, replace Linear with CONV. The Q6 hint is Z_bias = 0 and S_bias = S_input · S_weight.
Q10 (5 points): compare the two approaches
The last question is open-ended: compare the pros and cons of K-means and linear quantization in terms of accuracy, latency, hardware support, and so on.
Before writing, revisit the side-by-side table on Lecture 6 page 3, which compares storage and computation for the two. Then check it against the numbers you printed in Q3 and Q9: K-means needs a codebook to reconstruct floating-point values, while linear quantization can stay in integers along the whole path. Your answer should rest on results you ran yourself.
How to self-study the lab
- Run through the FP32 baseline cell before anything else. Make sure the Colab GPU, weight download, and CIFAR-10 download all work before you start on questions.
- Run the matching test after every function. Q1, Q4, and Q7 have ready-made checks. Don't move on until they pass; later questions amplify earlier mistakes.
- Work Q5 on paper first. Subtract
r_min = S(q_min − Z)fromr_max = S(q_max − Z)and the multiple-choice answers fall out. Once you understand it, Q5.3 is just translation into code. - Q9.2 and Q10 test understanding. When you're done, check your answers against the Lecture 6 slides.
One thing you can do tonight: open the notebook and run only the Setup and FP32 evaluation sections, and write down the accuracy and model size. Compare every quantization result against those two numbers, and you'll get a direct feel for how much you save and how much you lose.
Further reading
- Series entry point and course status: Reading MIT 6.5940 (series overview)
- The same workflow applied to pruning: Lab 1: Pruning
- 4-bit weight quantization for LLMs: Lab 4 + Lab 5: AWQ and an LLM on a laptop
References
- MIT 6.5940 Fall 2024 course page — Lab 2 release and due dates, Canvas submission, grading weights
- Lab 2 Colab notebook (Fall 2024) — every question number, point value, hint, and setup detail in this post
- Lec06-Quantization-II.pdf (Fall 2024) — the matching integer-inference formulas, per-channel quantization, calibration methods
- Lec05-Quantization-I.pdf (Fall 2024) — K-means and linear quantization basics
- MIT 6.5940 Fall 2026 course page — Fall 2026 Lab 2 schedule
- Han et al., Deep Compression (arXiv:1510.00149) — K-means quantization and centroid fine-tuning, cited in the notebook
- Jacob et al., Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (arXiv:1712.05877) — linear quantization, cited in the notebook
Loading...