Skip to content
All tags

#knowledge-distillation

1 posts

MIT 6.5940 L9 Knowledge Distillation: Teaching a Small Model Means Matching More Than Output Probabilities

Lecture 9 of MIT 6.5940 (Fall 2024) has five parts: what knowledge distillation (KD) is and why temperature matters; six things a student can match (logits, weights, features, gradients, sparsity patterns, relations); self and online distillation, which drop the fixed large teacher; KD for detection, segmentation, GANs, NLP, and LLMs; and Network Augmentation, built for tiny models. Raising the temperature from T=1 to T=10 moves the teacher's cat-vs-dog output from 0.982/0.017 to 0.599/0.401. That shift is where KD starts passing on dark knowledge.