Skip to content
All tags

#llm-inference

27 posts

CMU 11-868 HW6: Training Llama-2-7B with DeepSpeed ZeRO + LoRA, Serving with SGLang

The sixth 11-868 assignment is the first to set aside your homemade MiniTorch and use industry frameworks. Two problems, 50 points each. Problem 1: edit a DeepSpeed training script to turn on LoRA so Llama-2-7B can train on two 16GB V100s. Problem 2: fill in the TODOs of an SGLang inference script and tune parameters to make generation faster. The two problems want conflicting GPUs: SGLang doesn't support V100, so you need an L40S, A6000, or A100. The spring due date was April 13, and the assignment page publishes no grading tests.

CMU 11-868 L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention

11-868 spends two lectures on one question: how does an inference server handle many requests at once without wasting KV cache on the GPU? Lecture 22 (Lei Li) starts from SGLang's scheduling loop: ORCA's continuous batching, RadixAttention's radix tree for KV, sorting and routing by prefix hit rate, and hiding CPU scheduling behind GPU compute. Lecture 24 is given by vLLM author Woosuk Kwon: PagedAttention cuts KV cache into fixed-size blocks and virtualizes them with a block table, taking the batch on one A100 from 8 to 40. The second half covers how vLLM cuts CPU overhead, uses piecewise CUDA graphs, splits models across GPUs, and manages memory for hybrid architectures.

CMU 11-868 L19–L20 Model Quantization: GPTQ Saves Memory, and the Speedup Comes Along for the Ride

11-868 spends two lectures on quantization. L19 goes from BF16 and absmax/zero-point quantization to AdaQuant, ZeroQuant, and LLM.int8(). L20 is all GPTQ. GPTQ quantizes weights only: after each column is quantized, it uses second-order information to adjust the weights not yet quantized, and lazy batch updates plus a Cholesky trick let it scale to 175B. What it mainly saves is memory. Inference gets faster because single-batch decoding was already bottlenecked on reading weights; the amount of arithmetic does not shrink.

CMU 11-868 Serving at Scale: Prefill/Decode Disaggregation, KV Cache, and Heterogeneous Hardware

11-868 closes with five serving decks: Hao Zhang on DistServe, Vikram Mailthody on NVIDIA Dynamo, Junchen Jiang on LMCache, Mingxing Zhang on Mooncake and KTransformers, and Lei Li's map of serving frameworks. They share one question: once serving grows from one machine to a data center, where do the compute and the KV cache go? The argument runs in three steps. Measure goodput under latency SLOs instead of raw throughput. Put prefill and decode on separate GPUs. Let the KV cache spill from GPU memory into CPU memory, SSDs, and remote storage.

CMU 11-868 L08–L09: Choosing a Vocabulary, Emitting Tokens, and Why Speculative Decoding Is Fast

L08 goes from BPE to VOLT, a method co-authored by the lecturer Lei Li: vocabulary size has both a cost and a value, and VOLT finds the sweet spot by asking how much normalized entropy each added token removes, then solves it as an optimal transport problem. The second half covers LLaMA 3 growing its vocabulary from 32k to 128k and the cost of byte-level BPE splitting one Chinese character into three tokens. L09 moves from greedy decoding, sampling, and beam search to speculative decoding: a small model guesses N tokens and the big model checks them in one forward pass, because checking is cheaper than generating. It ends with EAGLE, which predicts final-layer features instead of tokens.

Reading MIT 6.5940: Song Han's Efficient AI Course Skipped a Year, So This Series Is Built on Fall 2024

MIT 6.5940 (TinyML and Efficient Deep Learning Computing) teaches how to make models smaller and faster so they fit on laptops, phones, and microcontrollers: pruning, quantization, NAS, distillation, LLM deployment, and distributed training. It was not offered in Fall 2025 because Song Han was on sabbatical, and the 2025 course URL returns 404. Fall 2026 is running, but as of 2026-09-30 only L1–L6 and Labs 0–1 are out. This series therefore follows Fall 2024, the latest complete edition: 23 slide decks, 23 videos, and Labs 0–5 are all public (A3). Fall 2026 is graded A2 and compared in every post.

MIT 6.5940 Lab 4 + Lab 5: Quantizing an LLM with AWQ, Then Running LLaMA2-7B on Your Own Laptop

Lab 4 is a Colab notebook that rebuilds AWQ step by step on OPT-1.3B: first see how badly 3-bit quantization hurts perplexity, then keep 1% of the salient channels in FP16 (Q1), then protect them by scaling instead and search for the best scale (Q2). Each question is worth 50 points, plus a bonus scored on perplexity. Lab 5 moves to C++: run 4-bit LLaMA2-7B-chat on your own computer with TinyChatEngine and write five versions of the W4A8 linear-layer kernel (loop unrolling, multithreading, SIMD, multithreading plus unrolling, and all combined), 20 points each, plus up to 20 bonus points for performance. This post covers the questions, points, setup, and limits for outside learners. No solutions.

MIT 6.5940 Lecture 13: LLM Deployment Through Quantization, Sparsity, and Serving

Lecture 13 sorts the ways to speed up LLM inference into three paths. Quantization: SmoothQuant moves the difficulty of activation outliers onto the weights to make W8A8 work, AWQ uses activation magnitudes to find the roughly 1% of weights that matter and protects them by scaling for W4A16, and QServe combines both into W4A8KV4. Sparsity: Wanda prunes weights by |W|·‖X‖, DejaVu and MoE use only part of the parameters per token, and SpAtten and H2O drop unimportant tokens. Serving: TTFT/TPOT metrics, PagedAttention, FlashAttention, speculative decoding, and continuous batching. On the slides, INT3 OPT-6.7B has a perplexity of 43.16 with RTN; scaling the salient channels by 2 brings it to 14.07.

NTU Hung-yi Lee ML 2026 Guide: Faster Generation, Part 1: Flash Attention and Why Moving Data Is the Bottleneck

In week 3 of ML 2026, Hung-yi Lee spends the first half of the inference lecture on one technique: Flash Attention. A GPU's execution units are fast, but their workbench (on-chip SRAM) is tiny, so data has to be carried to and from the warehouse (HBM). The carrying is the bottleneck. A naive softmax makes several round trips to the warehouse. Flash Attention assumes the current maximum is Amax, then multiplies by a correction factor when a larger value shows up. That lets it find the maximum, build the denominator, and compute the weighted sum in one pass, without ever materializing the attention weights. The output is identical to standard attention, no retraining is needed, and the cost is a little extra compute and a little brain strain.

NTU Hung-yi Lee ML 2026 Guide: HW3 LLM Fast Inference: Seven Speed-up Papers, Then Measuring Speculative Decoding, FlashAttention, and vLLM on a GPU

HW3 is 20 multiple-choice questions at 0.5 points each. No code is submitted; students answer a quiz on NTU COOL. The first 10 questions come from reading papers: four on speculative decoding (Leviathan et al., DeepMind's Speculative Sampling, Inference with Reference, SpecInfer) plus FlashAttention 1–3. The last 10 require filling TODOs in the Colab and analyzing the results: acceptance rate of a hand-written speculative decoder, speed-up curves for an assistant model vs n-gram under two prompt regimes, HBM reads and theoretical FlashAttention speed-up from T4 specs, vLLM prefix caching across turns and a cache invalidation test, and the effect of CPU offload on throughput. All questions are printed in both Mandarin and English in the homework PDF, so outsiders can do the whole thing; they just cannot get the official answers.

Hung-yi Lee ML 2026 HW8: Spending More Inference Compute — What Voting, Self-Certainty, and DeepConf Each Buy in Accuracy

HW8 involves no coding and no code submission. The TAs provide a finished Colab that runs Llama-3.2-1B-Instruct on the first 100 GSM8K questions and compares direct inference, Self-Consistency, Self-Certainty, and DeepConf (Confidence), sampling 16 reasoning traces per method. You read three papers, run the notebook, and answer 20 questions on NTU COOL: 18 about the papers and 2 about the Colab results. The prerequisite is Lecture 7 (Reasoning) of Lee's 2025 course. All questions are printed in hw8.pdf in Chinese and English, and the Colab is publicly downloadable. Only the COOL quiz and grades need an NTU account.

NTU Hung-yi Lee ML 2026 Guide: Faster Generation, Part 2: KV Cache Saves Time, Fills the Warehouse, and How to Slim It Down

KV Cache stores the keys and values already computed so decode does not recompute them, but every token costs memory. For Gemma 2 27B that is about 0.72MB per token, so an 80GB A100 holds only about 114k tokens. Hung-yi Lee then walks through ways to shrink it: let queries share keys and values (MQA, GQA), compress keys and values into one vector without ever decompressing (MLA), limit the attention span (Sliding Window, StreamingLLM), and drop keys and values nobody attends to (Scissorhands, H2O). He ends with cross-conversation prompt caching: it only hits when the prefix is identical, so a system prompt should put stable content first.

CME295 2026 Lecture 8, Written Ahead: Three Kinds of Noise, One Training Objective, and the Price of Parallel Decoding in Diffusion LLMs

The 2026 edition of CME295 gives diffusion LLMs a full lecture (Lecture 8, November 20), with five listed subtopics: continuous, discrete and masked diffusion, training, and inference. This pre-lecture edition works from the original papers (DDPM, D3PM, SEDD, MDLM, LLaDA and others): continuous noise costs about 64x the compute on text, and the [MASK] absorbing state won out; the training objective is a masked cross-entropy weighted by 1/t; the speed comes from filling several positions per step, yet LLaDA's main results decode one token per step, and Fast-dLLM needs a confidence threshold plus an approximate KV cache to reach up to a 27.6x speedup.

CME295 2026 Lecture 5 (Pre-Lecture Edition): LLM Systems, or How the Same Model Runs Several Times Faster

The 2026 edition of CME295 Lecture 5, "LLM systems" (October 30), lists seven topics: distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, FlashAttention, and hardware trade-offs. Written before the lecture, this post uses about 70 slides from the 2025 Lectures 3 and 4 plus the original papers to tie them into a single ledger: an H100 needs roughly 295 operations per byte moved to saturate its compute, while token-by-token generation does about 1 per byte of weights read, so most speedups are about moving less data.

aideep-dive

Learn Inference: Inference Engineering, Rebuilt with Dials You Can Turn

learn-inference.com is an unofficial interactive companion to Philip Kiely's Inference Engineering (256 pages, Baseten Books, free PDF). It follows the book's 8 chapters and 42 sections with rewritten explanations, turns intuition-heavy ideas like TTFT, P99, speculative decoding, and prefix-cache routing into slider-driven simulators, and ships a keyless JSON API and MCP server.

Self-Hosting Open-Source LLMs: Framework Choice, Hardware Math, and When It Beats APIs

Open-source models now match closed-source on coding benchmarks, but self-hosting isn't just picking a model — vLLM handles high-concurrency production serving, SGLang is 29% faster on prefix-heavy workloads, Ollama is the local dev default, and llama.cpp runs on the least hardware. A100 cloud rentals run ~$1.4-2.2/hr; self-hosting breaks even at roughly 100M tokens/month.

Self-Hosted Inference Overview: When Running Your Own Models Makes Sense

The key question in self-hosted inference isn't how fast the engine is — it's your GPU utilization. A fully saturated A100 costs ~$0.70 per million output tokens; at 10% utilization that becomes $7, more than most cloud APIs. This overview maps seven tools across three layers to help you decide which layer you need.

TensorRT-LLM: The Compile-for-Performance NVIDIA-Only LLM Inference Engine

TensorRT-LLM is NVIDIA's open-source LLM inference library (Apache 2.0). It offline-compiles model weights and compute graphs into optimized TensorRT engines, then serves them with custom CUDA kernels, in-flight batching, and multi-dimensional parallelism. The cost: NVIDIA GPUs only, compilation takes tens of minutes, and switching models or quantization means rebuilding.

How to Pick a Self-Hosted Inference Server: From Ollama to Xinference, Six Tools and Their Trade-Offs

Self-hosted inference servers fall into three layers: execution engine (llama.cpp), serving engine (vLLM, SGLang), and model management platform (Ollama, Xinference, Triton). Picking the right layer matters more than picking the right tool — ask where your bottleneck is before deciding where to add complexity.

Xinference: One Platform to Manage LLM, Embedding, Speech, and Image Models

Xinference wraps vLLM, SGLang, llama.cpp, Transformers, and MLX under a single management layer, using a Web UI and OpenAI-compatible API to manage LLMs, embedding, rerank, speech, and image models — suited for self-hosted deployments that need multiple model types to coexist. But the management layer's parsing logic also creates a larger attack surface than pure serving engines (CVE-2026-61539 is a case study).

CS336 Lecture 10: LLM Inference Is About Reading Weights and KV Cache Less Often

Lecture 10 separates prefill from decode: prefill parallelizes and is often compute-bound, while decode is sequential and commonly bandwidth-bound. GQA/MLA, quantization, speculative decoding, continuous batching, and PagedAttention reshape that cost.

aideep-dive

OpenRouter: One API Key for Multi-Model, Multi-Provider LLM Routing

OpenRouter exposes many models and inference endpoints through an OpenAI-compatible API, with provider ordering, failover, BYOK, and zero-data-retention controls in one routing policy.

aideep-dive

Sail Research: Trading Latency for Cost in Long-Horizon Agent Inference

Sail Research lets each inference request declare a completion window, scheduling patient background agents on cheaper capacity, while Sailboxes provide persistent long-running execution environments.

vLLM: The Default Choice for Self-Hosted Inference — and When It's Over-Engineering

vLLM is the de facto standard for self-hosted LLM inference (89,470 GitHub stars, verified 2026-08-21), built on managing the KV cache the way an OS manages paged memory. But the selection question isn't how fast it is — it's your GPU utilization. Using Red Hat's measured 793 output tokens/second, a fully saturated A100 costs roughly $0.70 per million output tokens; at 10% utilization that becomes $7, more than most cloud APIs.

aiguide

llama.cpp — From Pure C++ to an LLM Inference Engine on Consumer Hardware

llama.cpp is the most widely used local LLM inference engine, implemented in pure C/C++. It supports CPU, Metal, CUDA, Vulkan, and other backends, and uses the GGUF quantization format to run multi-billion-parameter models on consumer hardware.

aiguide

TurboQuant+ — Two-Stage Quantization to Compress KV Cache to 2-bit, Running 100B Models on a MacBook

TurboQuant+ is an open-source implementation of a Google Research ICLR 2026 paper that uses PolarQuant + QJL two-stage quantization to compress the KV cache by 3.8-6.4x, enabling consumer hardware to run larger models with longer contexts.

vLLM — From PagedAttention to a Production-Grade LLM Inference Engine

vLLM uses PagedAttention to eliminate KV cache memory waste, combining continuous batching and prefix caching to become the most widely adopted open-source LLM inference engine today.