🌏 中文版
"Training teaches a model what it knows. Inference is everything that happens afterward, every time somebody uses it, and it is where the bill actually lands." That is the opening line of Learn Inference.
The site is an interactive companion to Philip Kiely's Inference Engineering, a 256-page book from Baseten Books that you can download as a free PDF. It follows the book's 8 chapters and 42 sections, rewrites the explanations, and replaces the parts that are "easier to understand by turning a dial than by reading a paragraph" with simulators. If you call LLM APIs but can't quite explain why a self-hosted model takes 400ms to produce its first word and then streams 30 per second, this site is written for you.
What it is: an unofficial interactive edition
First, the relationship to the book. The About page is blunt: this is an independent project with no connection to Baseten or Philip Kiely, nobody there reviewed or signed off on it, and when a rewritten explanation or figure gets something wrong, "the mistake was made here, not in the book." The author is unnamed; the About page says the site won't tell you, but the per-chapter Ask AI will if you ask.
So its scope equals the book's, delivered as web pages plus simulators. On audience, Baseten's book page is candid: anyone can read the first twenty pages, but after that some familiarity with coding and CS concepts helps. The site inherits the same bar.
How it compares with other ways to learn inference:
| Resource | Strength | Weakness |
|---|---|---|
| The book (PDF) | Complete, carries the author's own judgment | Static; you imagine the numbers |
| Stanford CS336's inference lecture | Derives from first principles, has assignments | Research-leaning; light on autoscaling and multi-cloud operations |
| Engine docs like vLLM | Most precise on configuration | Only covers itself, not the trade-off landscape |
| Learn Inference | Whole landscape plus simulators, free, no signup | Unofficial rewrite; errors are the site's own |
Why inference deserves its own study
Section 0.1 opens with a distinction that's easy to miss: training is a project, with a budget and an end date; inference is an operation, with no end date, load set by other people, and cost that scales with your success. What the rest of the book keeps returning to is inference's own two phases: prefill processes the whole prompt at once and is compute bound; decode reads the entire model from memory for every token and is bandwidth bound.
Those phases map to two numbers users actually feel: TTFT (how long until the first word) and TPS (how many words per second after that). The first simulator on the home page, taken from section 1.4, gives you a slider for each and a Run button to watch the response stream. The point it wants you to see:
A response that starts instantly and trickles can feel faster than one that pauses and then dumps, even when the second finishes first.
Read that sentence and you'll forget it. Drag the slider once and you won't. That's the bet the whole site makes.
For chapter 5's quantization, speculative decoding, KV cache re-use, parallelism, and disaggregation, the chapter intro frames it this way: each trades precision, memory, complexity, or hardware for latency or throughput, and the chapter is about knowing which trade you're making.
The simulators are the point
Every figure in the book is an interactive component on the site. llms.txt explains why: what the figures teach is "the response to input, which does not survive being written down." A few worth visiting on purpose:
- 1.4, why the mean hides your worst requests: push the tail weight up and the mean barely moves while P99 runs away. The caption calls that gap "the one in a hundred users who thinks your product is broken."
- 2.4, the roofline: the diagonal is what memory bandwidth allows, the flat top is what the tensor cores allow. Decode at batch size 1 sits far left of the ridge on every GPU you can buy, which is why batching exists.
- 5.2, where speculative decoding stops paying: lower the acceptance rate or lengthen the draft and the speedup drops below 1, meaning you're paying for compute to go slower. The section text adds a caveat people skip: at high batch sizes compute is no longer idle, and speculation can reduce total throughput.
- 5.3, the routing problem under prefix caching: a shared prefix only needs processing once, but the cache lives on one replica. Without cache-aware routing, an eight-replica fleet finds it one time in eight, and most of the theoretical win quietly disappears.
- 7.2, what a cold start is made of: four stages, and why a warm pool matters most — it removes the one stage you don't control, getting GPUs from your cloud provider.
One detail made me trust the site more: every figure is labeled with where its numbers come from — "Illustrative numbers," "Constants from the book," or "Datasheet numbers." It doesn't pass off illustrations as measurements.
Chapter map
| Ch. | Topic | In one line |
|---|---|---|
| 0 | Inference | Runtime, infrastructure, tooling: three layers, none optional |
| 1 | Prerequisites | Write "fast enough" as numbers before touching a kernel |
| 2 | Models | From linear layers to transformers to diffusion; find the bottleneck |
| 3 | Hardware | Reading GPU spec sheets, Hopper through Rubin |
| 4 | Software | CUDA → PyTorch → vLLM / SGLang / TensorRT-LLM, NVIDIA Dynamo |
| 5 | Techniques | Quantization, speculative decoding, caching, parallelism, disaggregation |
| 6 | Modalities | VLMs, embeddings, ASR, TTS, image and video generation |
| 7 | Production | Containers, autoscaling, multi-cloud GPUs, zero-downtime deploys, client code |
There's also a glossary and a further reading list grouped by area, taken from the book's Appendix B.
The side built for agents
The site is far friendlier to agents than most documentation. The Developers page lists four ways in:
- Append
.mdto any page URL for Markdown, or sendAccept: text/markdown - The whole book as one file at
/llms-full.txt, indexed at/llms.txt - A read-only JSON API,
GET /api/v1/chapters, with no account or key, described by an OpenAPI 3.1 spec - An MCP server exposing two tools,
list_chaptersandget_chapter, with no authentication
Hooking it into Claude Code or Claude Desktop takes one config block:
{
"mcpServers": {
"learn-inference": {
"url": "https://learn-inference.com/api/mcp"
}
}
}
I called the MCP server's tools/list and both tools came back. Note that the MCP server and API expose the chapter index, not the full text; for content, use .md or llms-full.txt. The book's author also warns that it's about 60,000 tokens, so don't feed it to an agent in one chunk.
Every chapter has an Ask AI box that answers about the page you're on. The privacy page says your question travels through Vercel's AI Gateway along with the current page, it can only read pages on the site, and your IP is hashed before being used for rate limiting.
How to read it
By role:
- Application engineers (calling APIs, building RAG or agents): read chapters 0 and 1, then jump to 5.3 Caching and 7.5 Client code. Something to do tonight: measure TTFT and P99 for your service separately instead of looking only at average latency.
- People about to self-host models: go 2.4 → 3 → 4.3 → 5, playing with each simulator before reading the text. Then pick up our open-source LLM self-hosting guide to map the concepts onto real framework choices.
- Interview prep: each of chapter 5's five sections is a common ML system design topic, and the simulator captions work as one-sentence answers.
Limits
- It isn't the book: the explanations are rewritten, and the site says errors are its own. If you're citing a number or argument, check it against the book.
- Lots of illustrative numbers: most simulators are labeled "Illustrative numbers." Good for intuition, not for estimating your own latency or cost.
- Anonymous author: no byline means no traceable expertise behind it; for teaching material, that means doing more of your own verification.
- NVIDIA-centric: the hardware and software chapters focus on NVIDIA's ecosystem, with other accelerators getting an overview in 3.4 — a scope inherited from the book.
References
- Learn Inference — home page
- Learn Inference: About — relationship to the book, non-affiliation statement
- Learn Inference: Developers — JSON API, MCP server, error format, versioning
- Learn Inference: llms.txt — site-wide chapter index and note on simulators
- Learn Inference: Privacy — where Ask AI data goes
- Learn Inference: 5.2 Speculative decoding
- Inference Engineering (Baseten Books) — the book by Philip Kiely, free download
- Stanford CS336: Inference — related post on this site
- vLLM inference engine — related post on this site
- Open-source LLM self-hosting guide — related post on this site
Loading...