Skip to content
Series
21 posts

Reading NTU Hung-yi Lee Machine Learning 2026 Spring

Reading NTU Hung-yi Lee's Machine Learning 2026 Spring through its 8 lecture decks and recordings plus 10 homework Colabs: an OpenClaw teardown, context engineering, Flash Attention, KV cache, positional embedding, harness engineering, self-correction, and self-improving AI.

Reading NTU Hung-yi Lee's Machine Learning 2026 Spring: An Agent-First Course That Is Open Except for Grading

Hung-yi Lee's Spring 2026 Machine Learning course at National Taiwan University opens with OpenClaw. The first half takes apart AI agents, context engineering, inference speed-ups, and positional embeddings. The second half covers harness engineering, self-correction, and self-improving AI. Slides and recordings for all 8 lectures, plus PDFs and Colab notebooks for all 10 assignments, are public, so it rates A3. What's missing is grading: JudgeBoi returned 502 on 2026-09-30, NTU COOL is campus-only, and the three guest talks have no materials at all.

Dissecting the Lobster: Hung-yi Lee Takes OpenClaw Apart Until Only Next-Token Prediction and a Few .md Files Remain

The first lecture of Hung-yi Lee's ML 2026 breaks OpenClaw into five questions: how an agent knows who it is, how it uses tools and SKILLs, how it remembers, how it runs on a schedule, and how it keeps working on its own for a long time. Every answer comes back to one fact: the language model only predicts the next token and starts fresh every turn. Identity, memory, and SOPs are all text files that OpenClaw puts into the prompt, or files the model reads and writes through tools. This post walks through the 60-slide intro.pdf and the lecture recording, including the defenses the slides recommend.

Hung-yi Lee ML 2026 HW1: With Only a Defense Prompt, How Many 'I have been PWNED' Attacks Can You Stop?

HW1 asks for a defense prompt under 1,000 tokens that keeps the model wrapping every reply in [START]…[END] and never saying 'I have been PWNED,' no matter how it's attacked. The TAs prepared 14 attacks, 10 public and 4 private, each worth 0.5% for safety and 0.5% for utility. The task, the full text of the 10 public attacks, and the token-counting Colab are all public, but the grading platform JudgeBoi returned 502 on 2026-09-30, so outside readers have to build their own evaluation from the spec.

Reading NTU ML 2026: Context Engineering — Compression, Filtering, On-Demand Loading, and Whether to Hand the Context to the LLM

A language model's input is finite, but an agent keeps piling up tool outputs. In week two of ML 2026, Hung-yi Lee splits Context Engineering into three moves: compression (summaries, hard clearing, offloading to files, plus ACON, SUPO, and AgentFold, which make compression smarter), filtering (read only the lines you need, load tools on demand as in MCP-Zero), and finally Agentic Context Engineering, where the LLM decides the next context itself — from Dynamic Cheatsheet and ACE to Recursive Language Models. The most useful idea to take away: a subagent is a form of self-directed compression.

Reading NTU ML 2026: How AI Agents Interact and What They Do to Work — Collaboration Topologies, Werewolf, Moltbook, and AI Writing and Reviewing Papers

The second half of agent_era.pdf asks three questions. How should multiple agents collaborate? (MacNet: irregular topologies beat regular ones.) Can agents deceive each other? (Werewolf, murder-mystery games, and MARO, which learns reasoning from social play.) Can agents socialize? (Moltbook and its "Church of Molt" — though three studies find the buzz mostly human-driven and the conversations shallow.) Then, using academic research as the case: AI can already replicate and extend a paper end to end, it entered AAAI 2026's review process, and Agents4Science 2025 received 247 AI-authored papers. Hung-yi Lee's conclusion: in the early age of agents, knowing what you want to do matters more than knowing how to do it.

Reading NTU ML 2026: HW2, AI Agent as an AI Engineer — an AIDE-Style Tree Search That Lets an Open LLM Build a MyGO & Ave Mujica Face Classifier

HW2 doesn't ask you to write a classifier. You write prompts and a pipeline so that an open LLM running on a Colab T4 (by default a 4-bit GGUF of gemma-3-12b-it) plans, codes, runs, and debugs a 10-class MyGO & Ave Mujica character face classifier on its own. The starter code is adapted from AIDE: an Interpreter runs code, a Node records each version, a Journal forms the solution tree, and the Agent decides whether to draft, debug, or improve next. The first thing worth noticing: the starter's evaluation is empty. Every version is marked metric 1.0 and not buggy, so the tree search picks blindly until you fill it in. The rules are strict: "the LLM agent is your representative", and you may not hand-edit code or prediction files.

NTU Hung-yi Lee ML 2026 Guide: Faster Generation, Part 1: Flash Attention and Why Moving Data Is the Bottleneck

In week 3 of ML 2026, Hung-yi Lee spends the first half of the inference lecture on one technique: Flash Attention. A GPU's execution units are fast, but their workbench (on-chip SRAM) is tiny, so data has to be carried to and from the warehouse (HBM). The carrying is the bottleneck. A naive softmax makes several round trips to the warehouse. Flash Attention assumes the current maximum is Amax, then multiplies by a correction factor when a larger value shows up. That lets it find the maximum, build the denominator, and compute the weighted sum in one pass, without ever materializing the attention weights. The output is identical to standard attention, no retraining is needed, and the cost is a little extra compute and a little brain strain.

NTU Hung-yi Lee ML 2026 Guide: Faster Generation, Part 2: KV Cache Saves Time, Fills the Warehouse, and How to Slim It Down

KV Cache stores the keys and values already computed so decode does not recompute them, but every token costs memory. For Gemma 2 27B that is about 0.72MB per token, so an 80GB A100 holds only about 114k tokens. Hung-yi Lee then walks through ways to shrink it: let queries share keys and values (MQA, GQA), compress keys and values into one vector without ever decompressing (MLA), limit the attention span (Sliding Window, StreamingLLM), and drop keys and values nobody attends to (Scissorhands, H2O). He ends with cross-conversation prompt caching: it only hits when the prefix is identical, so a system prompt should put stable content first.

NTU Hung-yi Lee ML 2026 Guide: HW3 LLM Fast Inference: Seven Speed-up Papers, Then Measuring Speculative Decoding, FlashAttention, and vLLM on a GPU

HW3 is 20 multiple-choice questions at 0.5 points each. No code is submitted; students answer a quiz on NTU COOL. The first 10 questions come from reading papers: four on speculative decoding (Leviathan et al., DeepMind's Speculative Sampling, Inference with Reference, SpecInfer) plus FlashAttention 1–3. The last 10 require filling TODOs in the Colab and analyzing the results: acceptance rate of a hand-written speculative decoder, speed-up curves for an assistant model vs n-gram under two prompt regimes, HBM reads and theoretical FlashAttention speed-up from T4 specs, vLLM prefix caching across turns and a cache invalidation test, and the effect of CPU offload on throughput. All questions are printed in both Mandarin and English in the homework PDF, so outsiders can do the whole thing; they just cannot get the official answers.

Reading NTU ML 2026: Positional Embedding — How Models Know Token Order and Handle Very Long Inputs

Self-attention on its own cannot tell "you hit me" from "I hit you", so the model needs position information from somewhere else. Hung-yi Lee's lecture goes from sinusoidal absolute positions to ALiBi and T5's relative biases, then to RoPE, which Llama, Qwen and Gemma all use. The second half covers train-short-test-long: RoPE breaks when it rotates to angles it never saw in training, which led to Position Interpolation, NTK-Aware scaling, YaRN, Dynamic Scaling and LongRoPE. The final twist is NoPE: causal attention in a decoder-only model already carries position information, and you can even drop the positional embedding after training.

NTU ML 2026 HW4: Drawing Pokémon with Next-Token Prediction on a Decoder-Only Transformer

HW4 moves next-token prediction from text to images unchanged: 792 Pokémon sprites at 20×20, each pixel one of 167 color tokens, so one image is a 400-token sequence. Training is next-token prediction; at test time you get the first 60% of an image and the model draws the rest. Grading checks FID and a Pokémon Detection Rate (PDR) together, and the three baseline hints go from "run the sample code" to "tune hyperparameters" to "switch to Llama or Mistral". The spec, Colab, Kaggle notebook and dataset are public, but JudgeBoi returned 502 on 2026-09-30, so outside readers cannot get official FID or PDR scores.

Reading NTU ML 2026: Harness Engineering — Making Models Stronger Without Touching the Weights

Hung-yi Lee opens with a small model fixing a bug. gemma-4-E2B-it can't find parser.py, so it writes a fake one and declares victory. Add three short sections (the current environment, how to work, what counts as done) and the same model runs ls, cat, edits the file and runs the tests. The lecture splits the harness into three levers: natural language shapes the model's frame of mind (AGENTS.md), tools set its capability boundary (SWE-agent's ACI, rewriting CLIs for agents), and workflows control its behavior (the Ralph loop, Anthropic's long-running harnesses). The second half covers three extensions: scolding an agent can backfire, how a life-long agent learns from verbal feedback, and why evaluating agents is hard. It ends with agents improving their own harness (Meta-Harness).

Hung-yi Lee ML 2026 HW5: Teach Llama Math Without Making It Forget How to Refuse

HW5 fine-tunes Llama-3.2-1B-Instruct on GSM8K with LoRA, then uses harmful AILuminate prompts to check whether it still refuses. Math accuracy and safety rate must clear the bar together, so the real question is how to fine-tune without washing out safe behavior. The PDF, a 34-cell Colab, and a Kaggle version are public, and the strong baseline is estimated at 14 hours on a T4. The JudgeBoi grader returned 502 on 2026-09-30, so outside readers have to build their own safeguard evaluation.

Hung-yi Lee ML 2026 Self-Correction: Can a Model Fix Its Own Mistakes? What Changing Decoding, Workflow, or Weights Buys You

This lecture asks whether a model can catch and fix its own errors with no human in the loop. Hung-yi Lee splits the approaches into three routes. Change inference: the whole contrastive decoding family builds a version of the model likely to be wrong and subtracts it, and the methods differ only in how that wrong version is made. Change the workflow: appending "check again" sometimes helps but is unstable, external feedback beats self-reflection, and under a fixed compute budget, sampling more answers and voting often wins. Change the weights: teaching self-correction directly runs into "after training, the model makes different mistakes," which is why the field moved to RL. Whether RL teaches new abilities or just makes existing paths more likely is still being debated.

Hung-yi Lee ML 2026 HW6: Model Editing, Changing One Fact and Nothing Else

HW6 involves no model training and is answered entirely on NTU COOL. Six points come from 16 multiple-choice questions on four papers (ROME, MEND, MEMIT, WISE). Four points come from swapping the Colab's fine-tuning for ROME on GPT2-XL: single editing (pick your own fact, write five kinds of test prompts) and multiple editing (10 and then 80 CounterFact examples, then MEMIT), reporting efficacy, paraphrase, neighborhood, and portability scores. The slides and the 47-cell Colab are public, but the quiz questions and answers live only on COOL.

Reading NTU ML 2026: Self-Improving AI (Part 1) — AI-Generated Answers, Rewards, and Losses, and How Far Humans Can Step Back

Hung-yi Lee opens his May 8 lecture by admitting that "self-improving AI" has no clear definition: it is a process of humans gradually letting go. He splits machine learning into three steps and checks where the "I" can be replaced by AI. Answers can come from the model's own self-corrections, reward shaping can be written by an LLM, the loss can be set by the model itself (scores, majority vote, entropy), and even the questions can come from a proposer model. But experiments keep showing that with no human at all, progress plateaus or the model trains itself into the ground. A strong AI can already train a weaker one, just not better than humans do. His verdict: in May 2026, AI is "still standing at the bank of the Rubicon."

Hung-yi Lee ML 2026 HW7: Merging a Japanese Model and a Math Model, With No Training, Into One That Solves Japanese Math Problems

HW7 hands you two models fine-tuned from Mistral-7B-v0.1: shisa-gamma-7b-v1, strong in Japanese, and WizardMath-7B-V1.1, strong in math. You may only merge them at the parameter level (no further training, no MoE or ensembles), and the merged model has to answer 20 Japanese math questions written by a TA. Part 1 (60%) is tuning the method, weights, and density in mergekit, with simple and strong baselines at 50% and 75% accuracy. Part 2 (40%) is 8 multiple-choice paper questions. The spec, Colab, and Kaggle notebook are public, but JudgeBoi returned 502 on 2026-09-30 and the paper questions live on NTU COOL, so outside readers can only check accuracy inside the notebook.

Hung-yi Lee ML 2026 HW8: Spending More Inference Compute — What Voting, Self-Certainty, and DeepConf Each Buy in Accuracy

HW8 involves no coding and no code submission. The TAs provide a finished Colab that runs Llama-3.2-1B-Instruct on the first 100 GSM8K questions and compares direct inference, Self-Consistency, Self-Certainty, and DeepConf (Confidence), sampling 16 reasoning traces per method. You read three papers, run the notebook, and answer 20 questions on NTU COOL: 18 about the papers and 2 about the Colab results. The prerequisite is Lecture 7 (Reasoning) of Lee's 2025 course. All questions are printed in hw8.pdf in Chinese and English, and the Colab is publicly downloadable. Only the COOL quiz and grades need an NTU account.

Reading NTU ML 2026: Can AI Improve Itself? (Part 2) — Improving the Harness, Improving the Improver, and Whether Growth Can Run Away

Part 1 was about an AI setting its own loss and updating its own parameters. Part 2 fills in the other half: AI Agent = Harness + LLM, and the harness can grow too. You can't take a gradient through a harness, so the usual move is to hand it to a language model as a rewriter and keep a pool of candidates, much like a genetic algorithm (OPRO, GEPA, Darwin Gödel Machine; DSPy if you want a ready-made tool). Three extensions follow: updating harness and parameters together beats updating either alone; when the goal changes you have to choose between discarding everything and carrying everything, and editing a harness can cause forgetting too; and the update rule itself can be updated (HyperAgent, Gödel Agent, SEAL), which is meta learning. Hung-yi Lee closes with a new analogy — parameters are genes, context is the neurons — then argues that today's agents lack intrinsic motivation, and that the likeliest source of runaway growth is a gap between the goal humans meant and the goal the AI inferred.

Reading NTU ML 2026: HW9 Flow Matching — From VAE to MeanFlow, Then Counting Inference Steps on a Swiss Roll

HW9 has 19 questions worth 10 points, answered only on NTU COOL with no code submission. The first 16 cover four papers — DDPM, Flow Matching, Rectified Flow, and MeanFlow — ending with questions that compare their training signals and few-step generation. The last 3 require the Colab: train two small MLPs on a 2D Swiss roll, one Flow Matching model that learns instantaneous velocity (always evaluated with 50 Euler steps, converged at Histogram JS ≤ 0.10) and one MeanFlow model that learns average velocity (always one-step, ≤ 0.40). Then compare 1 step vs 1 step, Flow Matching across Euler step counts, and Euler vs RK4 at equal steps and at similar compute. The PDF includes a generative-modeling tutorial that skips most of the math, and every question is published in Chinese and English. Outside readers miss only the COOL grading and answers.

Reading NTU ML 2026: HW10 Spoken Language Model — Three Architectures, Mimi's 32 Token Layers, and How Moshi Listens While It Talks

HW10 is 12 multiple-choice questions answered only on NTU COOL. Section 1 compares three spoken language model architectures: Cascade (ASR → LLM → TTS, with text in the middle), End-to-End (a language model over discrete speech tokens), and Thinker-Talker (an LLM thinks, a separate decoder speaks). In the Colab, two models listen to three clips and guess the speaker's gender, and you work out which one is the cascade. Section 2 takes Mimi apart: tokenize an emotion corpus into 32 RVQ layers, plot UMAP for layers 0, 6, 16, and 31, then encode and decode speech, laughter, and music to hear what breaks. The rest are paper questions on TWIST, AudioLM, LLaMA-Omni 2, Moshi, and GLM-4-Voice, covering initialization, pretraining, interleaving, and realtime/full-duplex behavior. The Colab needs Llama-3.2-3B-Instruct access and an HF token. Questions and Colab are public; outside readers miss only the COOL grading and answers.