Skip to content
All tags

#quantization

16 posts

CMU 11-868 L19–L20 Model Quantization: GPTQ Saves Memory, and the Speedup Comes Along for the Ride

11-868 spends two lectures on quantization. L19 goes from BF16 and absmax/zero-point quantization to AdaQuant, ZeroQuant, and LLM.int8(). L20 is all GPTQ. GPTQ quantizes weights only: after each column is quantized, it uses second-order information to adjust the weights not yet quantized, and lazy batch updates plus a Cholesky trick let it scale to 175B. What it mainly saves is memory. Inference gets faster because single-batch decoding was already bottlenecked on reading weights; the amount of arithmetic does not shrink.

CMU 11-868 L23 Efficient Fine-Tuning for Large Models: LoRA, CIAT, and QLoRA

Lecture 23 of 11-868 treats fine-tuning as a memory problem. Full-parameter half-precision fine-tuning of LLaMA-8B needs about 80GB. LoRA brings that to about 33GB, and QLoRA, which stores the frozen weights in 4 bits, gets it to about 9.2GB. The lecture moves in three steps: train only two small low-rank matrices A and B (the slides credit CIAT as the first to do this); squeeze frozen weights to about 0.52 bytes per parameter with an NF4 lookup table and double quantization; and use a paged optimizer to push optimizer state to the CPU when the GPU is about to run out.

Reading MIT 6.5940: Song Han's Efficient AI Course Skipped a Year, So This Series Is Built on Fall 2024

MIT 6.5940 (TinyML and Efficient Deep Learning Computing) teaches how to make models smaller and faster so they fit on laptops, phones, and microcontrollers: pruning, quantization, NAS, distillation, LLM deployment, and distributed training. It was not offered in Fall 2025 because Song Han was on sabbatical, and the 2025 course URL returns 404. Fall 2026 is running, but as of 2026-09-30 only L1–L6 and Labs 0–1 are out. This series therefore follows Fall 2024, the latest complete edition: 23 slide decks, 23 videos, and Labs 0–5 are all public (A3). Fall 2026 is graded A2 and compared in every post.

MIT 6.5940 L18 Efficient Diffusion Models: Save on Steps, Resolution, and Compute per Step

Diffusion is slow because one large network runs dozens to thousands of times, starting from pure noise. Lecture 18 first covers DDPM, conditioning, latent diffusion, SDEdit, and DreamBooth, then attacks the cost three ways: fewer steps (DDIM skips steps, progressive distillation halves the step count each round), less compute per step (DC-AE compresses images 64x, recomputing only the edited 1.7% region cuts MACs 8.2x, SVDQuant runs FLUX in 4-bit), and more devices (DistriFusion is up to 6.1x faster on 8 A100s).

MIT 6.5940 Lab 2: Implementing K-means and Linear Quantization, Down to an Integer-Only VGG

Lab 2 is a Colab notebook with 10 questions worth 100 points, built around a VGG pretrained on CIFAR-10. The first 3 questions cover K-means quantization: write the quantizer, work out how many clusters n bits gives you, write the centroid update, then compare accuracy at 8, 4, and 2 bits before and after fine-tuning. The other 7 cover linear quantization: write q = round(r/S) + Z, derive the scale and zero-point formulas, do per-channel weight quantization and bias quantization, write integer versions of the fully connected and convolution layers, and finally convert the whole model to INT8 for inference. This post lays out the questions, points, setup, and limits for outside learners. No solutions.

MIT 6.5940 Lab 4 + Lab 5: Quantizing an LLM with AWQ, Then Running LLaMA2-7B on Your Own Laptop

Lab 4 is a Colab notebook that rebuilds AWQ step by step on OPT-1.3B: first see how badly 3-bit quantization hurts perplexity, then keep 1% of the salient channels in FP16 (Q1), then protect them by scaling instead and search for the best scale (Q2). Each question is worth 50 points, plus a bonus scored on perplexity. Lab 5 moves to C++: run 4-bit LLaMA2-7B-chat on your own computer with TinyChatEngine and write five versions of the W4A8 linear-layer kernel (loop unrolling, multithreading, SIMD, multithreading plus unrolling, and all combined), 20 points each, plus up to 20 bonus points for performance. This post covers the questions, points, setup, and limits for outside learners. No solutions.

MIT 6.5940 Lecture 13: LLM Deployment Through Quantization, Sparsity, and Serving

Lecture 13 sorts the ways to speed up LLM inference into three paths. Quantization: SmoothQuant moves the difficulty of activation outliers onto the weights to make W8A8 work, AWQ uses activation magnitudes to find the roughly 1% of weights that matter and protects them by scaling for W4A16, and QServe combines both into W4A8KV4. Sparsity: Wanda prunes weights by |W|·‖X‖, DejaVu and MoE use only part of the parameters per token, and SpAtten and H2O drop unimportant tokens. Serving: TTFT/TPOT metrics, PagedAttention, FlashAttention, speculative decoding, and continuous batching. On the slides, INT3 OPT-6.7B has a perplexity of 43.16 with RTN; scaling the salient channels by 2 brings it to 14.07.

MIT 6.5940 Lecture 5: Number Formats, K-Means Quantization, and Linear Quantization

MIT 6.5940 Lecture 5 starts from one fact: an 8-bit integer add uses 30x less energy than a 32-bit float add. It reviews the bit layouts of INT, fixed point, FP32/FP16/BF16, FP8, and FP4, then covers two quantization methods. K-means quantization saves storage only, since computation stays in floating point. Linear quantization, r = S(q − Z), turns matrix multiplication, fully connected layers, and convolutions into integer arithmetic.

MIT 6.5940 Lecture 6: Quantization II — PTQ Granularity and Clipping, QAT and STE, Binarization, Mixed Precision

Lecture 6 is about what to do when quantization costs you accuracy. First, without retraining: use finer scale granularity (per-channel, group, MX), clip outliers (EMA, calibration batches, MSE, KL), and round smarter (AdaRound). If that isn't enough, retrain: QAT keeps a full-precision copy of the weights, runs fake quantization in the forward pass, and uses the STE to pass gradients straight through. In the whitepaper table the slides cite, MobileNetV1 drops to 0.1% accuracy under per-tensor INT8 PTQ and recovers to 70.7% with per-channel QAT, against a 70.9% float baseline. The last two sections cover 1–2 bit binary and ternary networks, and HAQ, which uses reinforcement learning to assign a bit width to each layer.

NTU ADL 2025 TA Recitations: From PyTorch and Hugging Face to LoRA, Quantization, and vLLM Deployment

The ADL Fall 2025 course page schedules seven TA recitations: Dev Infra (PyTorch, debugging) → NLP project lifecycle → the underlying logic of NLP projects → LLM LoRA training → LLM basics, architecture, and MoE → LLM inference and evaluation → LLM deployment. All ten videos are older recordings by Yen-Ting Lin from 2023 and 2024, reused in Fall 2025. The course page's five slide links all return 404; files with the same names still open under the Fall 2024 path, and Deployment has a video only. The first three sessions walk through the Hugging Face data → model → demo loop that HW1 needs; the last four cover training, inference, and serving LLMs.

Self-Hosting Open-Source LLMs: Framework Choice, Hardware Math, and When It Beats APIs

Open-source models now match closed-source on coding benchmarks, but self-hosting isn't just picking a model — vLLM handles high-concurrency production serving, SGLang is 29% faster on prefix-heavy workloads, Ollama is the local dev default, and llama.cpp runs on the least hardware. A100 cloud rentals run ~$1.4-2.2/hr; self-hosting breaks even at roughly 100M tokens/month.

Quantization & Inference Optimization: Running a 70B Model on Your Laptop

A 70B model needs ~140GB VRAM in FP16, but 4-bit quantization shrinks it to ~35GB. With llama.cpp's partial CPU offloading, it can run on consumer hardware. GGUF naming conventions (Q4_K_M, Q5_K_S) tell you the precision-size tradeoff. KV cache is why long conversations slow down.

CS336 Lecture 10: LLM Inference Is About Reading Weights and KV Cache Less Often

Lecture 10 separates prefill from decode: prefill parallelizes and is often compute-bound, while decode is sequential and commonly bandwidth-bound. GQA/MLA, quantization, speculative decoding, continuous batching, and PagedAttention reshape that cost.

aiguide

llama.cpp — From Pure C++ to an LLM Inference Engine on Consumer Hardware

llama.cpp is the most widely used local LLM inference engine, implemented in pure C/C++. It supports CPU, Metal, CUDA, Vulkan, and other backends, and uses the GGUF quantization format to run multi-billion-parameter models on consumer hardware.

aiguide

TurboQuant+ — Two-Stage Quantization to Compress KV Cache to 2-bit, Running 100B Models on a MacBook

TurboQuant+ is an open-source implementation of a Google Research ICLR 2026 paper that uses PolarQuant + QJL two-stage quantization to compress the KV cache by 3.8-6.4x, enabling consumer hardware to run larger models with longer contexts.

aiguide

Small Models That Run on Phones: Choices and Constraints in 2026

The main on-device LLMs in 2026 are Gemma 3n, Qwen 3.5 Small, Llama 3.2, Phi-4-mini, Ministral 3, and SmolLM3. Sub-3B quantized models can hit 30-50 tokens/sec on phones with 8GB RAM, but RAM, thermal throttling, and context window remain hard constraints.