Skip to content
All tags

#kv-cache

14 posts

CMU 11-868 L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention

11-868 spends two lectures on one question: how does an inference server handle many requests at once without wasting KV cache on the GPU? Lecture 22 (Lei Li) starts from SGLang's scheduling loop: ORCA's continuous batching, RadixAttention's radix tree for KV, sorting and routing by prefix hit rate, and hiding CPU scheduling behind GPU compute. Lecture 24 is given by vLLM author Woosuk Kwon: PagedAttention cuts KV cache into fixed-size blocks and virtualizes them with a block table, taking the batch on one A100 from 8 to 40. The second half covers how vLLM cuts CPU overhead, uses piecewise CUDA graphs, splits models across GPUs, and manages memory for hybrid architectures.

CMU 11-868 Serving at Scale: Prefill/Decode Disaggregation, KV Cache, and Heterogeneous Hardware

11-868 closes with five serving decks: Hao Zhang on DistServe, Vikram Mailthody on NVIDIA Dynamo, Junchen Jiang on LMCache, Mingxing Zhang on Mooncake and KTransformers, and Lei Li's map of serving frameworks. They share one question: once serving grows from one machine to a data center, where do the compute and the KV cache go? The argument runs in three steps. Measure goodput under latency SLOs instead of raw throughput. Put prefill and decode on separate GPUs. Let the KV cache spill from GPU memory into CPU memory, SSDs, and remote storage.

MIT 6.5940 L15 Long-Context LLM: When Context Grows, the KV Cache Breaks First

Lecture 15 has four parts. Extending context: interpolating RoPE stretches LLaMA from 2k to 32k, and LongLoRA's shifted sparse attention makes long-context fine-tuning cheap. Evaluation: lost-in-the-middle, Needle-in-a-Haystack, and LongBench. Efficient attention: the KV cache grows linearly with length. StreamingLLM finds that the first few tokens act as attention sinks, and keeping them plus a recent window gives stable generation. DuoAttention keeps a full KV cache only for a few retrieval heads. Quest keeps the whole KV cache but reads only the most critical pages for each query. The last part moves beyond Transformers: Mamba replaces attention with a selective SSM, and Jamba mixes the two.

MIT 6.5940 L12 Transformer and LLM: The Architecture Seen Through an Efficiency Lens

In Lecture 12, 6.5940 switches from CNNs to Transformers. The lecture doesn't dwell on theory. It points to where memory and compute go. Attention is O(N²). If Llama-2-70B used MHA, its KV cache at batch 16 and length 4096 would take 160GB. GQA shrinks that 8x and MQA shrinks it 64x. MoE adds total parameters while keeping per-token compute flat. This post bridges into Lecture 13 on LLM deployment.

NTU Hung-yi Lee ML 2026 Guide: HW3 LLM Fast Inference: Seven Speed-up Papers, Then Measuring Speculative Decoding, FlashAttention, and vLLM on a GPU

HW3 is 20 multiple-choice questions at 0.5 points each. No code is submitted; students answer a quiz on NTU COOL. The first 10 questions come from reading papers: four on speculative decoding (Leviathan et al., DeepMind's Speculative Sampling, Inference with Reference, SpecInfer) plus FlashAttention 1–3. The last 10 require filling TODOs in the Colab and analyzing the results: acceptance rate of a hand-written speculative decoder, speed-up curves for an assistant model vs n-gram under two prompt regimes, HBM reads and theoretical FlashAttention speed-up from T4 specs, vLLM prefix caching across turns and a cache invalidation test, and the effect of CPU offload on throughput. All questions are printed in both Mandarin and English in the homework PDF, so outsiders can do the whole thing; they just cannot get the official answers.

NTU Hung-yi Lee ML 2026 Guide: Faster Generation, Part 2: KV Cache Saves Time, Fills the Warehouse, and How to Slim It Down

KV Cache stores the keys and values already computed so decode does not recompute them, but every token costs memory. For Gemma 2 27B that is about 0.72MB per token, so an 80GB A100 holds only about 114k tokens. Hung-yi Lee then walks through ways to shrink it: let queries share keys and values (MQA, GQA), compress keys and values into one vector without ever decompressing (MLA), limit the attention span (Sliding Window, StreamingLLM), and drop keys and values nobody attends to (Scissorhands, H2O). He ends with cross-conversation prompt caching: it only hits when the prefix is identical, so a system prompt should put stable content first.

CME295 2026 Lecture 5 (Pre-Lecture Edition): LLM Systems, or How the Same Model Runs Several Times Faster

The 2026 edition of CME295 Lecture 5, "LLM systems" (October 30), lists seven topics: distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, FlashAttention, and hardware trade-offs. Written before the lecture, this post uses about 70 slides from the 2025 Lectures 3 and 4 plus the original papers to tie them into a single ledger: an H100 needs roughly 295 operations per byte moved to saturate its compute, while token-by-token generation does about 1 per byte of weights read, so most speedups are about moving less data.

Reading CMU 11-768 L3: How Long-Context Agents Manage Memory — Hybrid Attention, RoPE Extension, Prompt Caching, and Compaction

An agent resends its whole history on every call, so five calls already add up to 80K input tokens; 1,500 OpenHands sessions averaged 78K tokens, 37% of them tool results. Neubig works on two layers: at the model layer, hybrid attention (many local layers, one global) plus length curricula make million-token context possible; at the harness layer, stable prefixes earn cache reads roughly ten times cheaper, and compaction that keeps anchors and externalizes evidence gets past the limit — evaluated by how the agent continues afterwards.

Harvard CS181 HW6 (Part 1): Decoding Autoregressive Models, KV Cache, and Speculative Decoding

HW6 Problem 4 (20 points) takes apart the cost of generating one token at a time in three questions: picking the most likely token at each step doesn't give the most likely sequence; recomputing every key at every step makes cost quadratic, and a KV cache brings it back to linear; speculative decoding lets a small model guess and a large model verify in one pass. It is all pencil-and-paper, and every question maps onto a real design choice in today's LLM inference systems.

OMP append-only context: Why sync conversation by byte-stable prefix? How Anthropic/DeepSeek KV cache gets protected

omp uses StablePrefix to freeze system prompt + tool specs, AppendOnlyLog for append-only messages, and digest-based longestStablePrefix algorithm. When prune/shake/steering rewrite history, only the tail after the divergence point is resent. This maximizes Anthropic/DeepSeek prompt cache hit rate, fixing the old issue where every turn forced ~40k token re-prefill on llama.cpp (issue #3406).

Quantization & Inference Optimization: Running a 70B Model on Your Laptop

A 70B model needs ~140GB VRAM in FP16, but 4-bit quantization shrinks it to ~35GB. With llama.cpp's partial CPU offloading, it can run on consumer hardware. GGUF naming conventions (Q4_K_M, Q5_K_S) tell you the precision-size tradeoff. KV cache is why long conversations slow down.

CS336 Lecture 10: LLM Inference Is About Reading Weights and KV Cache Less Often

Lecture 10 separates prefill from decode: prefill parallelizes and is often compute-bound, while decode is sequential and commonly bandwidth-bound. GQA/MLA, quantization, speculative decoding, continuous batching, and PagedAttention reshape that cost.

Context and Memory: Where Agents Actually Fail

Chroma tested 18 frontier models and all of them degrade as input grows — as a cliff, not a slope. Memory failures are usually retrieval failures in disguise. And the real cost of KV cache is bandwidth, not storage: every generated token reads the whole cache.

aiguide

TurboQuant+ — Two-Stage Quantization to Compress KV Cache to 2-bit, Running 100B Models on a MacBook

TurboQuant+ is an open-source implementation of a Google Research ICLR 2026 paper that uses PolarQuant + QJL two-stage quantization to compress the KV cache by 3.8-6.4x, enabling consumer hardware to run larger models with longer contexts.