Skip to content

Reading MIT 6.5940: Song Han's Efficient AI Course Skipped a Year, So This Series Is Built on Fall 2024

Sep 30, 20261 min
TL;DRMIT 6.5940 (TinyML and Efficient Deep Learning Computing) teaches how to make models smaller and faster so they fit on laptops, phones, and microcontrollers: pruning, quantization, NAS, distillation, LLM deployment, and distributed training. It was not offered in Fall 2025 because Song Han was on sabbatical, and the 2025 course URL returns 404. Fall 2026 is running, but as of 2026-09-30 only L1–L6 and Labs 0–1 are out. This series therefore follows Fall 2024, the latest complete edition: 23 slide decks, 23 videos, and Labs 0–5 are all public (A3). Fall 2026 is graded A2 and compared in every post.

🌏 中文版

MIT 6.5940 is a graduate course taught by Song Han. His slide covers list him as Associate Professor at MIT and Distinguished Scientist at NVIDIA. The course URL, efficientml.ai, redirects to MIT HAN Lab's course page. The course tackles a practical problem: models grow faster than hardware, so how do you compress and speed them up enough to run on a laptop, a phone, or a microcontroller with a few hundred KB of memory?

The Fall 2024 course description lists model compression, pruning, quantization, neural architecture search, distributed training, data/model parallelism, gradient compression, and on-device fine-tuning, plus acceleration techniques for LLMs and diffusion models. It also promises a concrete hands-on result: students deploy Llama2-7B on their own laptop.

This post is the entry point to the series. It answers four questions: what the course teaches, why this series uses Fall 2024 instead of the newest edition, what outside readers can actually get, and how to read it.

The hard facts

ItemFall 2024 (series backbone)Fall 2026 (in progress)
TitleTinyML and Efficient Deep Learning ComputingTinyML and Efficient AI Computing
Course page2024-fall-659402026-fall-65940
Prerequisites6.191 Computation Structures and 6.390 Intro to Machine Learning; students without them are de-registered in week two, with a petition optionSame two subjects; the page says no prerequisite waivers and no cross-registration
Lectures23, plus 3 final project presentation sessions (Dec 3, 5, 10)22 (the last is a Guest Lecture), plus 3 presentation sessions (Dec 3, 8, 10)
ExamsNone; the page says the class "does not have any tests or exams"Same
Videos23, on the MIT HAN Lab YouTube channel, entry point live.efficientml.aiUploaded as the term runs; L1–L6 as of 2026-09-30
Submission / discussionCanvas and Piazza (enrolled MIT students only)Same

The course page calls it a "PhD level course." It sets two goals: understand efficient deep learning techniques, and be able to deploy an LLM on your own laptop.

Why Fall 2024 is the backbone

This site's course guides normally follow the latest complete edition from 2025–2026. 6.5940 has no complete edition in that window:

  1. Fall 2025 was cancelled. The "Time" field on the Fall 2024 page now reads: "The course will not be offered in Fall 2025 due to Prof. Han is on sabbatical." Opening hanlab.mit.edu/courses/2025-fall-65940 returns HTTP 404. The "Previous Courses" list at the bottom of the page shows only three 6.5940 editions (2026, 2024, 2023) plus its 2022 predecessor, 6.S965.
  2. Fall 2026 is not finished. Its schedule runs to December 10. As of 2026-09-30, L1–L6 have slides and videos, and Labs 0 and 1 are out. From L7 on, the Slides and Video links are still empty.

So this series uses Fall 2024, the latest complete edition. Each post has a "Fall 2026 comparison" section that lists the matching new material and what changed. As later Fall 2026 lectures appear, they will be added to each post through an update log. The post order stays the same.

Openness: Fall 2024 is A3, Fall 2026 is A2

The grades follow the A0–A3 definitions in this site's global AI/CS course map.

Fall 2024: A3 (enough to self-study). Every lecture on the course page links to Dropbox slides and a YouTube video. Labs 0–4 are public Colab notebooks, Lab 5 is a public Google Drive folder, and the final project list is a public Google Doc. You get systematic material plus the assignment files, which meets A3.

A3 still has gaps. Know these before you start:

  • No official solutions. Labs are submitted on Canvas, so outside readers get no grading feedback and have to check their work against the slides.
  • The four "Chapter" slots are empty. Chapter I–IV on the schedule (Sep 11, Oct 16, Nov 11, Nov 20) are section markers with empty Slides and Video links.
  • L22 only has summary slides. L22 is titled "Course Summary + Quantum Machine Learning I," but its slide link is the 13-page Course-Summary.pdf, which has no quantum ML content. Quantum ML Part I exists only as video.
  • Final presentations were not recorded. None of the three presentation sessions has a link.

Fall 2026: A2 (partly open). Slides and videos appear as the term runs, currently through L6. The course page says it does not accept cross-registered students.

Grading and labs

Fall 2024 grading, from the course page:

ItemWeight
5 labs15% each, 75% total
Final project25% (Proposal 5% + Presentation and Final Report 20%)
Participation Bonus4% (end-of-term course survey)

Other rules: labs are individual, though discussing them is allowed if you name your collaborators. There are 6 penalty-free late days for the whole term. You must turn in at least 4 of the 5 labs to pass. On team size, the course page says groups of 4 or 5, while slide 89 of Lecture 1 says "group of 3-4," so the two official sources disagree. The report is 4 pages in the NeurIPS template, and slide 11 of the Course Summary deck also asks for a GitHub link to open-source the code. The project list was released on 2024-10-24.

Fall 2024 labs:

LabTopicLectures
Lab 0PyTorch tutorialL2
Lab 1PruningL3–L4
Lab 2QuantizationL5–L6
Lab 3Neural architecture searchL7–L8
Lab 4LLM compressionL13
Lab 5LLM deployment on laptopL13

Slide 88 of Lecture 1 lists the minimum hardware for Lab 5: macOS, Linux, or Windows; an x86 or ARM (Apple M1/M2) processor; 8 GB of memory; 5 GB of free storage.

Fall 2026 grading is not on the course page. It appears only on slide 88 of the Fall 2026 Lecture 1 deck: 5 labs at 14% each, a 30% final project (Proposal 5% + Presentation and Final Report 25%), and a Class Survey Bonus with no stated weight.

What changed from Fall 2024 to Fall 2026

This table lists only differences that the two course pages and slide decks confirm directly.

ItemFall 2024Fall 2026
L1–L6Introduction, Basics, Pruning I/II, Quantization I/IISame topics; the L2 Lecture Plan adds a CNN architecture review (AlexNet, VGG-16, ResNet-50, MobileNetV2)
L13Efficient LLM DeploymentLLM Quantization and Deployment
L17–L18GAN, Video, and Point Cloud; Diffusion ModelDiffusion Model Part I and Part II
End of termL22 Course Summary + Quantum ML I; L23 Quantum ML IIDec 1 Guest Lecture (speaker and topic not announced)
Lab 1PruningGPU Basics (titled "Lab1: Efficient AI Fundamentals" inside the zip)
Lab 2, Lab 4Quantization; LLM compressionCourse page: Quantization, Quantization. L1 slides: Pruning, Quantization
GradingLabs 15% × 5, project 25%, bonus 4%Labs 14% × 5, project 30%, survey bonus

The Lab 2 row needs a note. The Fall 2026 course page lists "Lab2: Quantization" and "Lab4: Quantization," so Quantization appears twice in the same list. Slides 87 and 89 of Lecture 1 say "Lab 2 — Pruning" instead. The two official sources contradict each other, and this series will not guess. It will update once Lab 2 is actually released. Until then, readers who want to practice pruning should use Fall 2024 Lab 1.

Fall 2026 Lab 1 can already be confirmed. Its README lists five parts: latency and MAC/FLOPs/I/O, the roofline model, a Gemma-3 decoder layer case study (prefill vs. decode), PyTorch Profiler and kernel fusion, and Flash Attention. It is worth 80 points plus 20 bonus points. Part 5 needs an A100 GPU on Colab. Fall 2024 had no GPU profiling lab like this, so the series gives it a separate supplementary post.

Course map: four chapters

The schedule uses four chapter markers to group the 23 lectures:

ChapterLecturesTopics
Chapter I: Efficient InferenceL3–L11Pruning, quantization, NAS, knowledge distillation, MCUNet, TinyEngine
Chapter II: Domain-Specific OptimizationL12–L18Transformers and LLMs, LLM deployment, post-training, long context, ViT, GAN/video/point cloud, diffusion
Chapter III: Efficient TrainingL19–L21Distributed training, on-device training and transfer learning
Chapter IV: Advanced TopicsL22–L23Course summary, quantum ML

L1–L2 come before Chapter I and supply the motivation and the measuring tools. Slides 6–7 of the Course Summary deck cut the material another way: Efficient Inference, Efficient Training, and Application-Specific Optimizations, placed on an Algorithm axis and a System axis. The course teaches algorithms and systems together, and that is its biggest difference from a typical deep learning course.

Series contents

The series has 25 posts (order 0–24). There are more posts than lectures because the pruning, quantization, and NAS labs are large enough to get their own posts.

OrderPostMaterial
0This post: series entryBoth course pages
1Why efficiency matters and how to measure model size and computeL1, L2, Lab 0
2Pruning I: granularity and criteriaL3
3Pruning II: per-layer ratios and hardware supportL4
4Lab 1: fine-grained and channel pruningF24 Lab 1
5Quantization I: number formats and basic quantizationL5
6Quantization II: PTQ, QAT, and mixed precisionL6
7Lab 2: k-means and linear quantizationF24 Lab 2
8NAS I: search space and search strategyL7
9NAS II: hardware-aware NASL8
10Lab 3: searching subnets under constraintsF24 Lab 3
11Knowledge distillationL9
12MCUNet: neural networks on microcontrollersL10
13TinyEngine and parallel computingL11
14Bridge: Transformers and LLMsL12
15LLM deploymentL13
16Labs 4 + 5: AWQ and an LLM on your laptopF24 Labs 4 and 5
17Fall 2026 supplement: Lab 1 GPU BasicsF26 Lab 1
18LLM post-trainingL14
19Long-context LLMsL15
20Efficient vision: ViT, GANs, video, point cloudsL16, L17
21Accelerating diffusionL18
22Distributed trainingL19, L20
23On-device trainingL21
24Course summary and quantum MLL22, L23

Three reading routes

The full route. Read orders 0–24 in sequence, with each lecture's slides and video, and do each lab when you reach its post. Fall 2024 released a lab every two or three lectures. At that pace the route takes about a semester.

LLM efficiency only. Read order 1 (the measuring tools) → 5 and 6 (quantization basics, needed for AWQ) → 14 → 15 → 16 → 17 → 19. You can skip pruning and NAS for now and come back to order 2 when you reach LLM sparsity.

TinyML and edge devices. Read order 1 → 2, 3, 4 → 5, 6, 7 → 8, 9, 10 → 12 → 13 → 23. This is Chapter I plus on-device training, focused on design under KB-scale memory limits.

Before you start

  • Python and PyTorch. Lab 0 is a PyTorch tutorial. If it feels hard, catch up first.
  • Backpropagation and CNNs. L2 only reviews terms and layer shapes; it does not teach how training works. If that is new to you, read the MIT 6.7960 guide or the CMU 11-785 guide first.
  • Basic computer architecture. The 6.191 prerequisite is a computer architecture course, and TinyEngine, SIMD, and the memory hierarchy come up later.
  • A Google account. Labs 0–4 run on Colab.

Further reading

These site series overlap with 6.5940. This series still covers the overlapping material in full; the links are for readers who want another angle:

Series navigation: Next: Why efficiency matters and how to measure model size and compute

References