A lecture-by-lecture reading of Stanford CME295: Transformers & Large Language Models, from the Transformer architecture to LLM evaluation and AI agents, and how it divides the ground with CS224N and CS336.
CME295 is a two-unit Stanford course with no homework; your grade is the midterm and the final, 50% each. The 2025 edition's nine lectures are fully public: videos, slides, and both exams with solutions. The 2026 edition rewrites the agent lecture around context compaction, harnesses, coding agents, and skills, and adds three full lectures on LLM systems, reinforcement learning, and Diffusion LLMs.
CME295 Lecture 1 threads a single sentence, "A cute teddy bear is reading.", through the whole class: split it into tokens, turn them into vectors, see why an RNN can't hold on to long sentences, then translate it into French with self-attention and an encoder-decoder. The 2026 edition drops the entire section on NLP tasks and evaluation metrics and opens instead with a timeline running from 2017 to the agent era.
CME295 Lecture 2 takes the original Transformer apart and refits it: position information moves from "added to the embedding" to RoPE's "rotate Q and K inside attention"; attention gets cheaper with sliding windows and MQA/GQA; models split into encoder-only, encoder-decoder, and decoder-only families; and the second half dissects BERT's MLM (15% of tokens) and NSP pretraining. The 2026 edition folds all of this into a single "Large Language Models" lecture, and BERT is no longer a syllabus item.
CME295 Lecture 3 defines an LLM as a decoder-only next-token predictor, uses MoE to explain why a huge model only touches part of its weights per token, and spends most of its time on the knobs you can turn at generation time: greedy, beam search, top-k, top-p, temperature, guided decoding, plus three prompting techniques (few-shot, chain of thought, self-consistency). The 2026 edition folds this lecture into Lecture 2, and the prompting half disappears from the syllabus.
CME295 Lecture 4 splits LLM training into two stages: pretraining on trillions of tokens (Llama 3 used 15 trillion), then SFT on thousands to millions of demonstrations so the model stops continuing text and starts answering. In between sits a map of memory savers (ZeRO, FlashAttention, mixed precision); the lecture closes with LoRA and QLoRA, which let people without big GPUs finetune, with QLoRA cutting VRAM by about 16x on a 65B model.
SFT only teaches a model to imitate good answers; it has no way to say which answers are unacceptable. CME295 Lecture 5 covers how to collect preference pairs, walks through the two steps of RLHF (a reward model trained on roughly 10,000 human labels, then PPO on roughly 100,000 examples), and ends with DPO, which folds the whole RL pipeline into a single supervised loss. The 2026 edition splits this lecture between Lecture 3 (training) and a new Lecture 4 (reinforcement learning).
CME295 Lecture 6 breaks reasoning models into three pieces: emit a reasoning chain before the answer, run RL on verifiable rewards like "is the answer correct," and use GRPO, which takes the group's average reward as the baseline instead of training a value model. RL alone took DeepSeek-R1-Zero from 15.6% to 71.0% pass@1 on AIME 2024, and distilling R1's traces into Qwen-32B beat running RL on the 32B model directly.
Lecture 7 of CME295 (2025) patches three LLM gaps: RAG fixes knowledge frozen at training time with a two-stage retrieve-then-rerank pipeline; tool calling fixes the inability to act by having a backend execute the function call the model writes; agents chain those calls with ReAct's observe-plan-act loop. The 2026 edition renames it AI Agents and adds context compaction, harness optimization, coding agents, and skills, the biggest rewrite in the course.
CME295 Lecture 8 starts from the fact that human rating is slow and expensive and BLEU/ROUGE can't recognize a paraphrase. It covers how LLM-as-a-Judge works, three biases (position, verbosity, self-enhancement) and six best practices, splits agent failures into tool prediction, tool execution and response generation, and closes with what MMLU, AIME, SWE-bench, HarmBench and τ-bench each measure, plus pass^k and Goodhart's law.
The last CME295 lecture packs 128 slides into three parts: an eight-picture recap of the quarter, how Transformers handle images (ViT and two ways to build a VLM), and masked diffusion LLMs that emit several tokens per step, followed by what comes next in research and applications. It is not on the exam; the 2026 edition turns diffusion LLMs into a lecture of their own and refocuses Lecture 9 on multimodality.
The 2026 edition of CME295 Lecture 5, "LLM systems" (October 30), lists seven topics: distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, FlashAttention, and hardware trade-offs. Written before the lecture, this post uses about 70 slides from the 2025 Lectures 3 and 4 plus the original papers to tie them into a single ledger: an H100 needs roughly 295 operations per byte moved to saturate its compute, while token-by-token generation does about 1 per byte of weights read, so most speedups are about moving less data.
The 2026 CME295 Lecture 4 (October 16) gives RL its own lecture, with seven syllabus items: mathematical conventions, reward design, policy gradients, limitations, PPO, GRPO, and on-policy distillation. This post walks the math ahead of class: start from ∇log π times a score. SFT uses a score of 1, PPO estimates it with a value model, GRPO uses the group mean, and on-policy distillation uses the teacher's per-token log-prob gap. In the Qwen3 report, starting from the same checkpoint, RL reached 67.6 on AIME'24 with 17,920 GPU hours, while on-policy distillation reached 74.4 with about 1/10 of that (1,800 hours).
The 2026 syllabus for CME295 Lecture 6 (November 6, 2026) lists seven topics. Tool calling, MCP and retrieval were already covered in the 2025 Lecture 7; the genuinely new ones are context compaction, harness optimization, coding agents, and skills/plugins. This pre-lecture edition explains those four using engineering posts from Anthropic and OpenAI, the MCP 2026-07-28 spec, and the Meta-Harness paper.
The 2026 edition of CME295 gives diffusion LLMs a full lecture (Lecture 8, November 20), with five listed subtopics: continuous, discrete and masked diffusion, training, and inference. This pre-lecture edition works from the original papers (DDPM, D3PM, SEDD, MDLM, LLaDA and others): continuous noise costs about 64x the compute on text, and the [MASK] absorbing state won out; the training objective is a masked cross-entropy weighted by 1/t; the speed comes from filling several positions per step, yet LLaDA's main results decode one token per step, and Fast-dLLM needs a confidence threshold plus an approximate KV cache to reach up to a 27.6x speedup.