Skip to content
All tags

#model-serving

9 posts

CMU 11-868 L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention

11-868 spends two lectures on one question: how does an inference server handle many requests at once without wasting KV cache on the GPU? Lecture 22 (Lei Li) starts from SGLang's scheduling loop: ORCA's continuous batching, RadixAttention's radix tree for KV, sorting and routing by prefix hit rate, and hiding CPU scheduling behind GPU compute. Lecture 24 is given by vLLM author Woosuk Kwon: PagedAttention cuts KV cache into fixed-size blocks and virtualizes them with a block table, taking the batch on one A100 from 8 to 40. The second half covers how vLLM cuts CPU overhead, uses piecewise CUDA graphs, splits models across GPUs, and manages memory for hybrid architectures.

CMU 11-868 Serving at Scale: Prefill/Decode Disaggregation, KV Cache, and Heterogeneous Hardware

11-868 closes with five serving decks: Hao Zhang on DistServe, Vikram Mailthody on NVIDIA Dynamo, Junchen Jiang on LMCache, Mingxing Zhang on Mooncake and KTransformers, and Lei Li's map of serving frameworks. They share one question: once serving grows from one machine to a data center, where do the compute and the KV cache go? The argument runs in three steps. Measure goodput under latency SLOs instead of raw throughput. Put prefill and decode on separate GPUs. Let the KV cache spill from GPU memory into CPU memory, SSDs, and remote storage.

MIT 6.5940 Lecture 13: LLM Deployment Through Quantization, Sparsity, and Serving

Lecture 13 sorts the ways to speed up LLM inference into three paths. Quantization: SmoothQuant moves the difficulty of activation outliers onto the weights to make W8A8 work, AWQ uses activation magnitudes to find the roughly 1% of weights that matter and protects them by scaling for W4A16, and QServe combines both into W4A8KV4. Sparsity: Wanda prunes weights by |W|·‖X‖, DejaVu and MoE use only part of the parameters per token, and SpAtten and H2O drop unimportant tokens. Serving: TTFT/TPOT metrics, PagedAttention, FlashAttention, speculative decoding, and continuous batching. On the slides, INT3 OPT-6.7B has a perplexity of 43.16 with RTN; scaling the salient channels by 2 brings it to 14.07.

aideep-dive

Learn Inference: Inference Engineering, Rebuilt with Dials You Can Turn

learn-inference.com is an unofficial interactive companion to Philip Kiely's Inference Engineering (256 pages, Baseten Books, free PDF). It follows the book's 8 chapters and 42 sections with rewritten explanations, turns intuition-heavy ideas like TTFT, P99, speculative decoding, and prefix-cache routing into slider-driven simulators, and ships a keyless JSON API and MCP server.

Xinference: One Platform to Manage LLM, Embedding, Speech, and Image Models

Xinference wraps vLLM, SGLang, llama.cpp, Transformers, and MLX under a single management layer, using a Web UI and OpenAI-compatible API to manage LLMs, embedding, rerank, speech, and image models — suited for self-hosted deployments that need multiple model types to coexist. But the management layer's parsing logic also creates a larger attack surface than pure serving engines (CVE-2026-61539 is a case study).

aideep-dive

Fireworks AI: From Serverless APIs to Custom Model Deployments

Fireworks AI puts open-weight model evaluation, dedicated GPU deployments, and LoRA customization behind one API surface. Serverless fits low-volume starts, On-demand fits sustained traffic and custom models, while reserved capacity adds enterprise capacity guarantees.

NVIDIA Triton Inference Server: Multi-Framework Models, Dynamic Batching, and Pipelines

Triton Inference Server serves TensorRT, ONNX, PyTorch, and other models through consistent HTTP and gRPC APIs. Its defining tools are the model repository, dynamic batching, instance groups, and ensembles—not LLM-specific KV-cache scheduling.

aideep-dive

Stanford CS25 V6: A Course Called Transformers United Whose First Two Talks Weren't About Transformers

CS25 is Stanford's 1-unit seminar where attendance is the only homework and anyone can audit. Of the nine talks in the Spring 2026 season, the three worth your time are Albert Gu on the inductive biases of SSMs vs Transformers, Charles Frye on serving inference across thousands of GPUs, and Victoria Lin on what native multimodality still hasn't solved.

vLLM — From PagedAttention to a Production-Grade LLM Inference Engine

vLLM uses PagedAttention to eliminate KV cache memory waste, combining continuous batching and prefix caching to become the most widely adopted open-source LLM inference engine today.