Skip to content
All tags

#vllm

12 posts

CMU 11-868 L22 and L24 LLM Serving: Scheduling, RadixAttention, and PagedAttention

11-868 spends two lectures on one question: how does an inference server handle many requests at once without wasting KV cache on the GPU? Lecture 22 (Lei Li) starts from SGLang's scheduling loop: ORCA's continuous batching, RadixAttention's radix tree for KV, sorting and routing by prefix hit rate, and hiding CPU scheduling behind GPU compute. Lecture 24 is given by vLLM author Woosuk Kwon: PagedAttention cuts KV cache into fixed-size blocks and virtualizes them with a block table, taking the batch on one A100 from 8 to 40. The second half covers how vLLM cuts CPU overhead, uses piecewise CUDA graphs, splits models across GPUs, and manages memory for hybrid architectures.

NTU ADL 2025 TA Recitations: From PyTorch and Hugging Face to LoRA, Quantization, and vLLM Deployment

The ADL Fall 2025 course page schedules seven TA recitations: Dev Infra (PyTorch, debugging) → NLP project lifecycle → the underlying logic of NLP projects → LLM LoRA training → LLM basics, architecture, and MoE → LLM inference and evaluation → LLM deployment. All ten videos are older recordings by Yen-Ting Lin from 2023 and 2024, reused in Fall 2025. The course page's five slide links all return 404; files with the same names still open under the Fall 2024 path, and Deployment has a video only. The first three sessions walk through the Hugging Face data → model → demo loop that HW1 needs; the last four cover training, inference, and serving LLMs.

NTU Hung-yi Lee ML 2026 Guide: HW3 LLM Fast Inference: Seven Speed-up Papers, Then Measuring Speculative Decoding, FlashAttention, and vLLM on a GPU

HW3 is 20 multiple-choice questions at 0.5 points each. No code is submitted; students answer a quiz on NTU COOL. The first 10 questions come from reading papers: four on speculative decoding (Leviathan et al., DeepMind's Speculative Sampling, Inference with Reference, SpecInfer) plus FlashAttention 1–3. The last 10 require filling TODOs in the Colab and analyzing the results: acceptance rate of a hand-written speculative decoder, speed-up curves for an assistant model vs n-gram under two prompt regimes, HBM reads and theoretical FlashAttention speed-up from T4 specs, vLLM prefix caching across turns and a cache invalidation test, and the effect of CPU offload on throughput. All questions are printed in both Mandarin and English in the homework PDF, so outsiders can do the whole thing; they just cannot get the official answers.

Self-Hosting Open-Source LLMs: Framework Choice, Hardware Math, and When It Beats APIs

Open-source models now match closed-source on coding benchmarks, but self-hosting isn't just picking a model — vLLM handles high-concurrency production serving, SGLang is 29% faster on prefix-heavy workloads, Ollama is the local dev default, and llama.cpp runs on the least hardware. A100 cloud rentals run ~$1.4-2.2/hr; self-hosting breaks even at roughly 100M tokens/month.

Self-Hosted Inference Overview: When Running Your Own Models Makes Sense

The key question in self-hosted inference isn't how fast the engine is — it's your GPU utilization. A fully saturated A100 costs ~$0.70 per million output tokens; at 10% utilization that becomes $7, more than most cloud APIs. This overview maps seven tools across three layers to help you decide which layer you need.

How to Pick a Self-Hosted Inference Server: From Ollama to Xinference, Six Tools and Their Trade-Offs

Self-hosted inference servers fall into three layers: execution engine (llama.cpp), serving engine (vLLM, SGLang), and model management platform (Ollama, Xinference, Triton). Picking the right layer matters more than picking the right tool — ask where your bottleneck is before deciding where to add complexity.

Xinference: One Platform to Manage LLM, Embedding, Speech, and Image Models

Xinference wraps vLLM, SGLang, llama.cpp, Transformers, and MLX under a single management layer, using a Web UI and OpenAI-compatible API to manage LLMs, embedding, rerank, speech, and image models — suited for self-hosted deployments that need multiple model types to coexist. But the management layer's parsing logic also creates a larger attack surface than pure serving engines (CVE-2026-61539 is a case study).

CS336 Lecture 10: LLM Inference Is About Reading Weights and KV Cache Less Often

Lecture 10 separates prefill from decode: prefill parallelizes and is often compute-bound, while decode is sequential and commonly bandwidth-bound. GQA/MLA, quantization, speculative decoding, continuous batching, and PagedAttention reshape that cost.

vLLM: The Default Choice for Self-Hosted Inference — and When It's Over-Engineering

vLLM is the de facto standard for self-hosted LLM inference (89,470 GitHub stars, verified 2026-08-21), built on managing the KV cache the way an OS manages paged memory. But the selection question isn't how fast it is — it's your GPU utilization. Using Red Hat's measured 793 output tokens/second, a fully saturated A100 costs roughly $0.70 per million output tokens; at 10% utilization that becomes $7, more than most cloud APIs.

aiproject

2026 Q1 Open-Source LLM Landscape: From Frontier Models to On-Device, a Complete Survey

2026 Q1 saw a full-blown open-source model explosion: on the LLM front, GLM-5, Kimi K2.5, and Qwen3.5 caught up with closed-source models; Embedding and Reranker are dominated by Qwen3 and BGE; speech has Voxtral TTS and Whisper V3; image has FLUX.2; and video has Wan 2.2 rivaling Sora. This is the complete navigation map.

OpenClaw's 60 Providers: A Category Map, and What Actually Bites When You Attach a Local Model

The official provider directory now lists 60 entries. The most common failure when attaching a local model is writing Ollama's base URL with /v1 — that breaks tool calling, and the model starts emitting raw tool-call JSON as plain text.

vLLM — From PagedAttention to a Production-Grade LLM Inference Engine

vLLM uses PagedAttention to eliminate KV cache memory waste, combining continuous batching and prefix caching to become the most widely adopted open-source LLM inference engine today.