Skip to content
All tags

#cme295

14 posts

CME295 Lecture 7: Agentic LLMs, or Letting the Model Look Things Up, Call Functions, and Run Its Own Loop

Lecture 7 of CME295 (2025) patches three LLM gaps: RAG fixes knowledge frozen at training time with a two-stage retrieve-then-rerank pipeline; tool calling fixes the inability to act by having a backend execute the function call the model writes; agents chain those calls with ReAct's observe-plan-act loop. The 2026 edition renames it AI Agents and adds context compaction, harness optimization, coding agents, and skills, the biggest rewrite in the course.

CME295 2026 Lecture 6 Preview: AI Agents, from Calling Tools to Managing Context and Tuning the Harness

The 2026 syllabus for CME295 Lecture 6 (November 6, 2026) lists seven topics. Tool calling, MCP and retrieval were already covered in the 2025 Lecture 7; the genuinely new ones are context compaction, harness optimization, coding agents, and skills/plugins. This pre-lecture edition explains those four using engineering posts from Anthropic and OpenAI, the MCP 2026-07-28 spec, and the Meta-Harness paper.

CME295 Lecture 9: Transformers Leave Text Behind, and LLMs Stop Writing Left to Right

The last CME295 lecture packs 128 slides into three parts: an eight-picture recap of the quarter, how Transformers handle images (ViT and two ways to build a VLM), and masked diffusion LLMs that emit several tokens per step, followed by what comes next in research and applications. It is not on the exam; the 2026 edition turns diffusion LLMs into a lecture of their own and refocuses Lecture 9 on multimodality.

CME295 2026 Lecture 8, Written Ahead: Three Kinds of Noise, One Training Objective, and the Price of Parallel Decoding in Diffusion LLMs

The 2026 edition of CME295 gives diffusion LLMs a full lecture (Lecture 8, November 20), with five listed subtopics: continuous, discrete and masked diffusion, training, and inference. This pre-lecture edition works from the original papers (DDPM, D3PM, SEDD, MDLM, LLaDA and others): continuous noise costs about 64x the compute on text, and the [MASK] absorbing state won out; the training objective is a masked cross-entropy weighted by 1/t; the speed comes from filling several positions per step, yet LLaDA's main results decode one token per step, and Fast-dLLM needs a confidence threshold plus an approximate KV cache to reach up to a 27.6x speedup.

CME295 Lecture 3: The Knobs You Turn When an LLM Generates, from Temperature and Top-p to Chain of Thought

CME295 Lecture 3 defines an LLM as a decoder-only next-token predictor, uses MoE to explain why a huge model only touches part of its weights per token, and spends most of its time on the knobs you can turn at generation time: greedy, beam search, top-k, top-p, temperature, guided decoding, plus three prompting techniques (few-shot, chain of thought, self-consistency). The 2026 edition folds this lecture into Lecture 2, and the prompting half disappears from the syllabus.

CME295 Lecture 8: Using LLMs to Judge LLMs, and the Three Biases to Guard Against

CME295 Lecture 8 starts from the fact that human rating is slow and expensive and BLEU/ROUGE can't recognize a paraphrase. It covers how LLM-as-a-Judge works, three biases (position, verbosity, self-enhancement) and six best practices, splits agent failures into tool prediction, tool execution and response generation, and closes with what MMLU, AIME, SWE-bench, HarmBench and τ-bench each measure, plus pass^k and Goodhart's law.

CME295 Lecture 6: How Reasoning Models Learn to Think Longer, and What GRPO Drops from PPO

CME295 Lecture 6 breaks reasoning models into three pieces: emit a reasoning chain before the answer, run RL on verifiable rewards like "is the answer correct," and use GRPO, which takes the group's average reward as the baseline instead of training a value model. RL alone took DeepSeek-R1-Zero from 15.6% to 71.0% pass@1 on AIME 2024, and distilling R1's traces into Qwen-32B beat running RL on the 32B model directly.

CME295 2026 Lecture 5 (Pre-Lecture Edition): LLM Systems, or How the Same Model Runs Several Times Faster

The 2026 edition of CME295 Lecture 5, "LLM systems" (October 30), lists seven topics: distributed training, inference optimizations, KV caching, speculative decoding, efficient kernels, FlashAttention, and hardware trade-offs. Written before the lecture, this post uses about 70 slides from the 2025 Lectures 3 and 4 plus the original papers to tie them into a single ledger: an H100 needs roughly 295 operations per byte moved to saturate its compute, while token-by-token generation does about 1 per byte of weights read, so most speedups are about moving less data.

CME295 Lecture 4: The Bill for Training an LLM, and Where Pretraining, SFT, and LoRA Spend It

CME295 Lecture 4 splits LLM training into two stages: pretraining on trillions of tokens (Llama 3 used 15 trillion), then SFT on thousands to millions of demonstrations so the model stops continuing text and starts answering. In between sits a map of memory savers (ZeRO, FlashAttention, mixed precision); the lecture closes with LoRA and QLoRA, which let people without big GPUs finetune, with QLoRA cutting VRAM by about 16x on a 65B model.

CME295 Lecture 5: SFT Can't Teach "Don't Answer Like That", So RLHF and DPO Add the Negative Signal

SFT only teaches a model to imitate good answers; it has no way to say which answers are unacceptable. CME295 Lecture 5 covers how to collect preference pairs, walks through the two steps of RLHF (a reward model trained on roughly 10,000 human labels, then PPO on roughly 100,000 examples), and ends with DPO, which folds the whole RL pipeline into a single supervised loss. The 2026 edition splits this lecture between Lecture 3 (training) and a new Lecture 4 (reinforcement learning).

CME295 2026 Lecture 4 (Pre-Lecture Edition): SFT, PPO, GRPO, and On-Policy Distillation Are One Policy Gradient

The 2026 CME295 Lecture 4 (October 16) gives RL its own lecture, with seven syllabus items: mathematical conventions, reward design, policy gradients, limitations, PPO, GRPO, and on-policy distillation. This post walks the math ahead of class: start from ∇log π times a score. SFT uses a score of 1, PPO estimates it with a value model, GRPO uses the group mean, and on-policy distillation uses the teacher's per-token log-prob gap. In the Qwen3 report, starting from the same checkpoint, RL reached 67.6 on AIME'24 with 17,920 GPU hours, while on-policy distillation reached 74.4 with about 1/10 of that (1,800 hours).

CME295 Lecture 1: From Tokens to Transformer, or How One Sentence Gets Translated into Another Language

CME295 Lecture 1 threads a single sentence, "A cute teddy bear is reading.", through the whole class: split it into tokens, turn them into vectors, see why an RNN can't hold on to long sentences, then translate it into French with self-attention and an encoder-decoder. The 2026 edition drops the entire section on NLP tasks and evaluation metrics and opens instead with a timeline running from 2017 to the agent era.

CME295 Lecture 2: How One Transformer Grew into BERT, GPT, and a Zoo of Attention Variants

CME295 Lecture 2 takes the original Transformer apart and refits it: position information moves from "added to the embedding" to RoPE's "rotate Q and K inside attention"; attention gets cheaper with sliding windows and MQA/GQA; models split into encoder-only, encoder-decoder, and decoder-only families; and the second half dissects BERT's MLM (15% of tokens) and NSP pretraining. The 2026 edition folds all of this into a single "Large Language Models" lecture, and BERT is no longer a syllabus item.

Reading Stanford CME295: Two Units, No Homework, Nine Lectures from Transformers to AI Agents

CME295 is a two-unit Stanford course with no homework; your grade is the midterm and the final, 50% each. The 2025 edition's nine lectures are fully public: videos, slides, and both exams with solutions. The 2026 edition rewrites the agent lecture around context compaction, harnesses, coding agents, and skills, and adds three full lectures on LLM systems, reinforcement learning, and Diffusion LLMs.