An RNN translator has to squeeze the whole source sentence into one vector, and long sentences don't fit. Attention lets the decoder look back at every input position each time it produces a word, score each one, and take a weighted sum. Call the scorer the query, the thing being scored the key, and the thing being averaged the value, and you have dot-product attention. The Transformer goes one step further: the input attends to itself, recurrence disappears, and in return you get parallelism and a constant path length. The price is that position information has to be added back by hand.
A static word vector gives "apple" one embedding, whether it means the fruit or the company. The BERT lecture in ADL Fall 2025 starts from that polysemy problem. TagLM feeds language-model features into a tagger, ELMo builds contextual embeddings from a deep bidirectional LSTM, and BERT swaps the LSTM for a Transformer, pre-trained with Masked LM and Next Sentence Prediction; downstream, you add a classifier or tagger on the top layer and fine-tune. The optional BERT Variants slides go one ring further out: Transformer-XL for longer context, XLNet's permutation LM to get both AR and AE benefits, RoBERTa's better data and training recipe, SpanBERT's span masking, and mBERT and XLM for many languages. This is the direct prerequisite for HW1, which uses bert-base-chinese for extractive QA.
Big data is not big annotated data. The last ADL lecture asks how to learn good representations without labels, and answers: find the latent factors that control the data. An auto-encoder squeezes the input into a short code and reconstructs it. The denoising version adds noise or masks 15% of tokens first, which is exactly the idea behind BERT's masked LM. A VAE forces the code to follow a distribution, so you can sample from it to generate. Dual learning lets paired tasks, such as translation and back-translation or understanding and generation, act as feedback for each other. Self-supervised learning has two camps: self-prediction (hide part, guess it back) and contrastive learning (pull similar pairs together, push dissimilar ones apart). CLIP runs contrastive learning on 400 million image-text pairs, making zero-shot image classification possible, and DALL·E 2 uses CLIP's representations to generate images. Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.
Dialogue systems split into chit-chat and task-oriented. Task-oriented systems were traditionally built from four modules: language understanding (LU) turns a sentence into domain, intent, and slots; dialogue state tracking (DST) accumulates the user's goal; the dialogue policy picks the next system action; and NLG turns that action back into a sentence. An LLM can act out all four steps by itself, but it cannot actually make the booking, so it needs external tools. LaMDA learns to call a search engine, calculator, and translator. BlenderBot 2.0 adds internet search and long-term memory. WebGPT learns to drive a browser from human demonstrations, a reward model, and PPO. Toolformer has the model generate and filter its own tool-use training data. The lecture ends with evaluation: automatic metrics, four kinds of human evaluation, and LLM-Eval. ADL Fall 2025 has only videos for this lecture, so the Fall 2024 slides fill in.
Applied Deep Learning (ADL) Fall 2025, taught by Yun-Nung (Vivian) Chen in NTU's CSIE department, is a deep learning course built around NLP. It runs from neural network basics through Transformers, BERT, pretraining and prompting, post-training, LoRA, RAG, generation and evaluation, alignment issues, and language agents. Lectures L0–L11 come with slide PDFs and segmented videos, and the playlist adds videos for L12–L14. On the assignment side only the HW1 spec is public; HW2, HW3, and the final project have explainer videos only. That makes it A2.
Lecture 10 of ADL Fall 2025 sorts the problems of pretrained models into four groups, each paired with a goal: bias with fairness, toxicity with safety, hallucination with factuality, and finally alignment. The slides argue that bias can enter at any stage of the ML pipeline, that safeguards belong at four layers (data, input, training, output), that hallucination can be checked atomic fact by atomic fact, and that over-optimizing a reward model produces familiar symptoms: verbosity, excessive apologies, over-refusal. The final project announced that week is called Jailbreaking Olympics, but all that is public is the titles and one-line descriptions of two videos.
HW1 in ADL Fall 2025 gives a question and four Chinese paragraphs. The model first picks the relevant paragraph (paragraph selection, framed as four-way multiple choice), then marks the answer's start and end inside it (span selection), scored by Exact Match. The spec slides point you straight at Hugging Face's run_swag_no_trainer.py and run_qa_no_trainer.py. The simple baseline uses bert-base-chinese, length 512, effective batch size 2, and learning rate 3e-5, and both stages together take under three hours on an 8GB RTX 3070. The Kaggle leaderboard closed 9/29, and code plus report were due 10/1 on NTU COOL. Outside readers cannot get the Kaggle data or grading, but the task design, baseline settings, and five report questions are all usable for practice.
Lecture 11 of ADL Fall 2025 builds on the EMNLP 2024 Language Agents tutorial. It defines an agent as an entity that perceives and acts, then names what is new about language agents: reasoning itself counts as an internal action. The lecture is organized around three concepts. Reasoning covers CoT and ReAct; memory covers Generative Agents and its recency / importance / relevance retrieval; planning goes from greedy reactive planning to tree search and world models. It closes with multi-agent systems in three steps: initialization, orchestration, and team optimization.
The first self-study deck of ADL Fall 2025 describes machine learning as finding a function from data, and deep learning as a production line of simple functions where the machine learns what every station does. It uses speech and vision to contrast deep and shallow models, credits big data and GPUs for the post-2010 breakthroughs, and uses the universality theorem to ask why networks should be deep rather than fat. The most practical part comes last: the output domain decides the learning task, and the architecture should fit the properties of the input domain.
ADL Fall 2025's NN Basics and Backpropagation decks break model training into three questions. What is the model? Layers of neurons, each computing z = Wa + b and then a nonlinearity. What makes a function good? A smaller loss. How do we pick the best one? Gradient descent, in practice mini-batch SGD. Backpropagation computes gradients for millions of parameters efficiently: the forward pass stores each layer's output, the backward pass sends an error signal δ back from the output layer, and multiplying the two gives each weight's gradient.
Lecture 9 of ADL Fall 2025 answers two questions. The model gives you a probability distribution at every step, so how do you pick a word from it? And once you have a sentence, how do you judge it? The slides start with teacher forcing and exposure bias to show the gap between training and generation, then compare greedy, beam search, sampling, top-k, and nucleus sampling, and file temperature and the penalties under 'control' rather than decoding algorithms. The evaluation half covers BLEU, ROUGE, perplexity, and LLM-Eval, then explains why you would use RL to optimize whole-sentence quality directly.
When an LLM is too big to fine-tune in full, the LLM Adaptation slides of NTU ADL Fall 2025 offer three ways to change only a small part of it: insert small Adapter modules into the Transformer, represent the weight update with low-rank matrices (LoRA), or learn only a prefix or soft prompt (prompt tuning). The slides conclude that no single method fits every task. For HW2, the only public information is its title, "LLM Tuning and Prompt Tuning for Classical Chinese Translation"; the data, baseline, and grading have no written spec.
A pre-trained model can continue text, but that does not mean it follows instructions. The Post-Training slides of NTU ADL Fall 2025 fix this in two steps. Instruction tuning (FLAN, T0) teaches the model to read task descriptions. RLHF then pulls its outputs toward human preference. Three limits of instruction tuning connect the two steps, a reward model and pairwise comparisons solve two practical RL problems, and InstructGPT's SFT → reward model → PPO pipeline ties it all together. ChatGPT runs the same pipeline on multi-turn dialogue.
Lecture 6 of ADL Fall 2025 sorts pre-trained models into three families: encoders (the BERT family, bidirectional context), decoders (the GPT series, good at generation), and encoder-decoders (BART and T5, pre-trained with denoising). It then names two practical obstacles of the pre-trained-model era: downstream labeled data is scarce, and models keep growing until one copy per task no longer fits. The slides' answer is prompt learning. GPT-3's in-context learning shows a model can do a task without updating parameters; hand-written hard prompts (template plus verbalizer, LM-BFF) then give way to soft prompts optimized as vectors (P-Tuning, Prefix-Tuning, Prompt Tuning); and Liu et al.'s prompting typology closes the lecture.
LLMs cannot memorize long-tail facts, their knowledge goes stale, and they cannot see private documents. The RAG slides of NTU ADL Fall 2025 open with an LLM hallucinating about the lecturer herself, then split RAG into indexing, retrieval, and generation: sparse (TF-IDF, BM25) and dense (DPR, Contriever) retrieval, how dense retrievers are trained, pre- and post-retrieval techniques including pointwise and pairwise reranking. A closing roadmap organizes RAG, RETRO, FLARE, Search-R1, and others by what, how, and when to retrieve. For HW3, only the title is public: "Retriever & Reranker Training for RAG".
The Reasoning lecture of NTU ADL Fall 2025 has no public slides. It exists only as five videos in the course playlist: 12.1 What is Reasoning?, 12.2 Short CoT, 12.3 Test-Time Scaling, 12.4 Learning to Reason (imitating others), and 12.5 RL for Reasoning (evolving reasoning through exploration). This post lays out that route from the video titles alone, then pairs it with the CoT, ReAct, and 'reasoning enlarges the action space' pages of the previous Language Agents deck. Technical detail is left to the site's CS224N and CME295 reasoning posts.
The first real lecture of ADL Fall 2025 has one through-line: a language model predicts the next word. It starts from one-hot vectors and co-occurrence matrices, moves through the zero-probability problem of n-grams and the smoothing that a neural LM gets for free, and ends at the RNN LM, which folds all previous words into a hidden state. BPTT and vanishing/exploding gradients are the training cost, and LSTM and GRU patch it with gating. The lecture closes by splitting applications into sequence input versus sequence output, which separates tagging from encoder-decoder models.
The ADL Fall 2025 course page schedules seven TA recitations: Dev Infra (PyTorch, debugging) → NLP project lifecycle → the underlying logic of NLP projects → LLM LoRA training → LLM basics, architecture, and MoE → LLM inference and evaluation → LLM deployment. All ten videos are older recordings by Yen-Ting Lin from 2023 and 2024, reused in Fall 2025. The course page's five slide links all return 404; files with the same names still open under the Fall 2024 path, and Deployment has a video only. The first three sessions walk through the Hugging Face data → model → demo loop that HW1 needs; the last four cover training, inference, and serving LLMs.
With whole words as units, any unseen word becomes UNK. With single characters, meaning is hard to reassemble. This 22-page ADL deck explains the mainstream compromise, subwords, and the most common way to build them, BPE. The core is a tiny corpus of 4 words and 16 occurrences: start from characters, merge the most frequent adjacent pair each round, and after 9 merges you have units like newest</w> and low</w>, which then segment the unseen words lowest and powest. It ends with a GPT-3 tokenizer screenshot where the Chinese version of a sentence takes more than twice as many tokens as the English.
Lectures 7 and 8 of Machine Learning Techniques open the aggregation part of the course. T7 sorts ways of combining hypotheses into uniform, linear, and any blending (stacking), shows with a few lines of algebra that uniform blending reduces variance, and then uses the bootstrap to create diverse g_t from the single dataset you have: that is bagging. T8 reinterprets the bootstrap as example weighting, then deliberately up-weights the examples the previous hypothesis got wrong so the next one is forced to differ, and votes with α_t = ln √((1−ε_t)/ε_t): that is AdaBoost. Practice with Fall 2024 HW6 Q4 and Q9, plus HW7's bootstrap and AdaBoost proofs and a 500-round AdaBoost-Stump experiment on madelon. There are no official solutions.
Hsuan-Tien Lin's Machine Learning Foundations (16 lectures) and Machine Learning Techniques (16 lectures) are two Mandarin-taught MOOCs. All 130 YouTube videos and 32 slide decks are free. The MOOCs alone are A2: since August 2025, free Coursera accounts can only view the first module, so the exercises sit behind a paywall. Add the Fall 2024 course page, which publishes HW0–HW7 and the final project spec, and you reach A3, minus the grading chain: no official solutions, Gradescope and NTU COOL are enrolled-only, and the Kaggle competition returns 404. Fall 2026 is running now as a flipped classroom; slides through week 4, hw0, and hw1 are public.
Lectures 9–11 of Machine Learning Techniques tie three models together with one thread: trees plus aggregation. T9 treats a decision tree as conditional aggregation and covers C&RT's binary branching, Gini and regression impurity, pruning, categorical features, and surrogate branches. T10 applies bagging to fully grown trees; add random subspaces and random projections and you get a random forest, with free OOB validation and permutation-based feature importance. T11 re-derives AdaBoost as steepest descent in function space on the exponential error, then swaps in squared error to get GBDT, which fits regressions to residuals. Practice with the impurity and gradient boosting proofs in Fall 2024 HW7. There are no official solutions.
Foundations Lecture 4 first shows that learning is impossible: from D alone, any guess outside D can be called wrong. That is No Free Lunch. It then reframes the question with marbles in a bin. If the data is drawn independently from one distribution, Hoeffding's inequality says the in-sample error E_in is probably close to the true error E_out. Checking one fixed h is only verification. Once the algorithm chooses among M hypotheses, a union bound charges 2M exp(−2ε²N). Conclusion: with a finite hypothesis set and small E_in, learning is feasible. What to do when M is infinite is the next lecture's job.
Techniques T16 re-sorts the whole course into three families: how to exploit features (kernels, aggregation, extraction, low-dimensional compression), how to optimize (gradients, equivalent problems, multiple steps), and how to fight overfitting (regularization, validation). It then uses four KDD Cup–winning models to show how the pieces combine in practice. The MOOC was recorded in 2016 and its deep learning stops at pre-training. The Fall 2024 on-campus course filled the gap with 302u (the ReLU family, Xavier/He initialization), 303u (momentum, RMSProp, Adam), a 2020 keynote deck, mlmai.ics, and 1126, an 11-model summary. The Fall 2026 versions of these files are scheduled for week 16 and currently return 404.
Fall 2024 had six Foundations assignments. HW0 is 20 multiple-choice math prerequisite questions. HW1–HW5 each have 12 problems plus a bonus: Q1–4 are auto-graded, Q5–12 are graded by TAs, the programming problems use rcv1, cpusmall, and mnist from the LIBSVM datasets site, and HW5 uses LIBLINEAR. HW1 and HW2 each include a problem where you argue with a ChatGPT-style answer. Fall 2026 has released hw0 and hw1: hw1 is now 16 multiple-choice problems with 4 secretly chosen for TA grading, and the data is the course's own hw1_train.dat. The new policy allows AI tools and vibe coding, but AI-generated code needs block-by-block comments in your own words. Neither semester publishes official solutions.
Lecture 5 of Machine Learning Techniques rewrites the soft-margin SVM in unconstrained form: ½wᵀw plus C times the total hinge error. That is an L2-regularized model, and a larger C means weaker regularization. The hinge error and logistic regression's cross-entropy are both convex upper bounds of the 0/1 error, so the SVM approximates L2-regularized logistic regression. For probability outputs, you can use Platt's two-level learning, running logistic regression on top of SVM scores, or use the representer theorem to do kernel logistic regression directly. Lecture 6 uses the same theorem to get the closed form β = (λI + K)⁻¹y for kernel ridge regression, but β is dense; switching to the ε-insensitive tube error gives SVR with sparse coefficients. Fall 2026 does not schedule these two lectures.
Lecture 3 of Machine Learning Techniques merges "feature transform + inner product" into a single kernel function K(x, x′). Training and prediction in the dual SVM only need K, so d̃ can be infinite: the Gaussian kernel corresponds to an infinite-dimensional transform. Lecture 4 admits the SVM can still overfit and introduces violations ξₙ and a parameter C, giving the soft-margin SVM. Its dual differs from the hard-margin one in exactly one way: αₙ gets an upper bound C. The value of αₙ sorts the data into non-SVs, free SVs, and bounded SVs, and the fraction #SV/N upper-bounds the leave-one-out error, a cheap way to rule out dangerous (C, γ).
The first three lectures of Machine Learning Foundations define machine learning as a flow chart: an unknown target function f generates data D, and an algorithm A picks g from a hypothesis set H, hoping g ≈ f. The simplest H (the perceptron) and A (PLA) then show the chart in action. On linearly separable data, PLA makes at most R²/ρ² updates; on non-separable data, use pocket instead. Lecture 3 sorts learning problems along four axes: output, label, protocol, and input. Foundations mostly deals with batch, supervised binary classification or regression on concrete features. Practice with Fall 2024 HW1 and Fall 2026 hw1.
Lecture 11 of ML Foundations compares PLA, linear regression, and logistic regression on the same score s = wᵀx. The three differ only in their error functions, and scaled cross-entropy upper-bounds the 0/1 error, so both regressions can do classification. The lecture then turns logistic regression into SGD by computing the gradient on one random example, and builds multiclass classifiers from binary ones with OVA and OVO. Lecture 12 uses a feature transform Φ to turn a circular boundary into a line in Z-space. The price is that computation and d_vc both grow with the dimension, so the advice is: try a linear model first. Practice problems are in Fall 2024 HW4.
Lecture 1 of Machine Learning Techniques turns "which separating line is best?" into an optimization problem. Once you fix the scale so that min yₙ(wᵀxₙ+b) = 1, maximizing the margin is the same as minimizing ½wᵀw, which is a standard QP. Lecture 2 uses Lagrange duality to trade a QP with d̃+1 variables for one with N variables and N+1 constraints, then uses the KKT conditions to recover (b, w) from α. Only the points with αₙ > 0, the support vectors, affect the answer. The dual still contains the inner product zₙᵀzₘ, so the dependence on dimension is not really gone until the kernel lecture.
Linear regression writes squared error as (1/N)‖Xw − y‖², sets the gradient to zero, and gets w_LIN = X†y in one step. The hat matrix H = XX† projects y onto the column space of X, which shows that on average E_out − E_in ≈ 2(d+1)/N. Logistic regression estimates P(+1|x) with θ(wᵀx); maximum likelihood turns into the cross-entropy error ln(1 + exp(−y wᵀx)). It has no closed-form solution, so you walk downhill along −∇E_in step by step. That is gradient descent.
Lectures 12 and 13 of Machine Learning Techniques open the third part, distilling hidden features. T12 starts from a linear combination of perceptrons: two layers can build AND and OR but not XOR, and one more layer fixes that, which is the multi-layer perceptron. It then replaces sign with tanh, derives backprop, and covers non-convex optimization, d_vc = O(VD), weight elimination, and early stopping. T13 discusses the challenges of deep networks, uses autoencoders as information-preserving encodings for layer-wise pre-training, treats denoising as regularization, and proves that the optimal linear autoencoder is spanned by the top eigenvectors of XᵀX, which is PCA. The videos date from 2016; modern deep learning is covered by the Fall 2024 302u/303u slides. Practice: Fall 2024 HW7 Q4, Q9, and bonus Q13.
Lecture 13 of ML Foundations defines overfitting as 'lower E_in but higher E_out' and uses experiments to find four causes: too little data, stochastic noise, an overly complex target (deterministic noise), and excessive model power. Lecture 14's remedy is regularization. It rewrites 'step back to H₂' as the constraint ‖w‖² ≤ C, then uses a Lagrange multiplier to turn it into minimizing E_in + (λ/N)wᵀw, which is weight decay. Back in VC theory, regularization shrinks the effective VC dimension d_EFF, and L1 buys sparse solutions. Practice problems: Fall 2024 HW4 Q8–9 and HW5 Q1, Q5–6, Q10.
Techniques T14 reinterprets the Gaussian SVM as a linear vote over distance-based similarities, which gives the RBF network. Too many centers overfit, so k-means picks a few prototypes, and k-means itself is alternating optimization. T15 starts from the Netflix ratings data: one-hot encode user IDs, feed them into a linear network with the tanh removed, and you get matrix factorization R ≈ VᵀW, learned by alternating least squares or SGD. The lecture closes with a map of extraction models: boosting, neural nets, RBF networks, matrix factorization, and k-NN. These two lectures exist only as MOOC material. Neither the Fall 2024 nor the Fall 2026 schedule covers them, and no public homework problem does either.
The Techniques half of Fall 2024 has two homework sets and a final project, and all three PDFs are public. HW6 covers kernels, soft-margin SVM, and aggregation; its programming part uses LIBSVM on the 3-vs-7 subproblem of mnist.scale to count support vectors, compute margins, and run 128 validation rounds. HW7 covers bootstrap, impurity, AdaBoost, gradient boosting, and neural networks; its programming part is a 500-round AdaBoost-Stump on madelon. The final project is a fictional baseball league, HTMLB: predict home-team wins across two Kaggle stages and write an English report of at most seven pages that compares at least four methods. There are no official solutions. On 2026-09-30 both Kaggle pages returned 404 without login, so outside readers probably cannot get the HTMLB data and should reproduce the same splits on a public dataset instead.
The Hoeffding guarantee from L4 carries an M, the number of hypotheses. Perceptrons have infinitely many lines, so M blows up. L5 stops counting hypotheses and counts how many ○× patterns (dichotomies) they can produce on N data points instead; the maximum is the growth function m_H(N). 2D perceptrons produce at most 14 patterns on 4 points, fewer than 2⁴ = 16, so 4 is their break point. L6, marked optional by the course, proves that any break point caps m_H(N) by a polynomial, which is what makes the VC bound work.
Lecture 15 of ML Foundations tackles model selection. Selecting by E_in overfits, and selecting by E_test is cheating. The compromise is to carve a validation set out of the training data, select by E_val, then retrain on all the data. The validation size K is a dilemma, with K = N/5 as the rule of thumb. Leave-one-out is almost unbiased but expensive and unstable, so in practice you use 5-fold or 10-fold. Lecture 16 closes with three principles, Occam's razor, sampling bias, and data snooping, and a 'Power of Three' recap: three related fields, three bounds, three linear models, three tools. Practice problems are in Fall 2024 HW5.
L7 names the largest non-break point the VC dimension d_VC, proves that d-dimensional perceptrons have d_VC = d + 1, and rewrites the VC bound as E_out ≤ E_in + a model-complexity penalty, so both too large and too small a d_VC hurt. Theory asks for N ≈ 10,000·d_VC examples; in practice 10·d_VC is often enough. L8 swaps the fixed target function for a distribution P(y|x) and shows the VC theory still holds under noise. The error measure should come from the application: a CIA fingerprint check that penalizes admitting an intruder 1000 times more can be reduced to plain classification by copying examples.
The second half of agent_era.pdf asks three questions. How should multiple agents collaborate? (MacNet: irregular topologies beat regular ones.) Can agents deceive each other? (Werewolf, murder-mystery games, and MARO, which learns reasoning from social play.) Can agents socialize? (Moltbook and its "Church of Molt" — though three studies find the buzz mostly human-driven and the conversations shallow.) Then, using academic research as the case: AI can already replicate and extend a paper end to end, it entered AAAI 2026's review process, and Agents4Science 2025 received 247 AI-authored papers. Hung-yi Lee's conclusion: in the early age of agents, knowing what you want to do matters more than knowing how to do it.
A language model's input is finite, but an agent keeps piling up tool outputs. In week two of ML 2026, Hung-yi Lee splits Context Engineering into three moves: compression (summaries, hard clearing, offloading to files, plus ACON, SUPO, and AgentFold, which make compression smarter), filtering (read only the lines you need, load tools on demand as in MCP-Zero), and finally Agentic Context Engineering, where the LLM decides the next context itself — from Dynamic Cheatsheet and ACE to Recursive Language Models. The most useful idea to take away: a subagent is a form of self-directed compression.
Hung-yi Lee's Spring 2026 Machine Learning course at National Taiwan University opens with OpenClaw. The first half takes apart AI agents, context engineering, inference speed-ups, and positional embeddings. The second half covers harness engineering, self-correction, and self-improving AI. Slides and recordings for all 8 lectures, plus PDFs and Colab notebooks for all 10 assignments, are public, so it rates A3. What's missing is grading: JudgeBoi returned 502 on 2026-09-30, NTU COOL is campus-only, and the three guest talks have no materials at all.
In week 3 of ML 2026, Hung-yi Lee spends the first half of the inference lecture on one technique: Flash Attention. A GPU's execution units are fast, but their workbench (on-chip SRAM) is tiny, so data has to be carried to and from the warehouse (HBM). The carrying is the bottleneck. A naive softmax makes several round trips to the warehouse. Flash Attention assumes the current maximum is Amax, then multiplies by a correction factor when a larger value shows up. That lets it find the maximum, build the denominator, and compute the weighted sum in one pass, without ever materializing the attention weights. The output is identical to standard attention, no retraining is needed, and the cost is a little extra compute and a little brain strain.
HW10 is 12 multiple-choice questions answered only on NTU COOL. Section 1 compares three spoken language model architectures: Cascade (ASR → LLM → TTS, with text in the middle), End-to-End (a language model over discrete speech tokens), and Thinker-Talker (an LLM thinks, a separate decoder speaks). In the Colab, two models listen to three clips and guess the speaker's gender, and you work out which one is the cascade. Section 2 takes Mimi apart: tokenize an emotion corpus into 32 RVQ layers, plot UMAP for layers 0, 6, 16, and 31, then encode and decode speech, laughter, and music to hear what breaks. The rest are paper questions on TWIST, AudioLM, LLaMA-Omni 2, Moshi, and GLM-4-Voice, covering initialization, pretraining, interleaving, and realtime/full-duplex behavior. The Colab needs Llama-3.2-3B-Instruct access and an HF token. Questions and Colab are public; outside readers miss only the COOL grading and answers.
HW2 doesn't ask you to write a classifier. You write prompts and a pipeline so that an open LLM running on a Colab T4 (by default a 4-bit GGUF of gemma-3-12b-it) plans, codes, runs, and debugs a 10-class MyGO & Ave Mujica character face classifier on its own. The starter code is adapted from AIDE: an Interpreter runs code, a Node records each version, a Journal forms the solution tree, and the Agent decides whether to draft, debug, or improve next. The first thing worth noticing: the starter's evaluation is empty. Every version is marked metric 1.0 and not buggy, so the tree search picks blindly until you fill it in. The rules are strict: "the LLM agent is your representative", and you may not hand-edit code or prediction files.
HW3 is 20 multiple-choice questions at 0.5 points each. No code is submitted; students answer a quiz on NTU COOL. The first 10 questions come from reading papers: four on speculative decoding (Leviathan et al., DeepMind's Speculative Sampling, Inference with Reference, SpecInfer) plus FlashAttention 1–3. The last 10 require filling TODOs in the Colab and analyzing the results: acceptance rate of a hand-written speculative decoder, speed-up curves for an assistant model vs n-gram under two prompt regimes, HBM reads and theoretical FlashAttention speed-up from T4 specs, vLLM prefix caching across turns and a cache invalidation test, and the effect of CPU offload on throughput. All questions are printed in both Mandarin and English in the homework PDF, so outsiders can do the whole thing; they just cannot get the official answers.
HW7 hands you two models fine-tuned from Mistral-7B-v0.1: shisa-gamma-7b-v1, strong in Japanese, and WizardMath-7B-V1.1, strong in math. You may only merge them at the parameter level (no further training, no MoE or ensembles), and the merged model has to answer 20 Japanese math questions written by a TA. Part 1 (60%) is tuning the method, weights, and density in mergekit, with simple and strong baselines at 50% and 75% accuracy. Part 2 (40%) is 8 multiple-choice paper questions. The spec, Colab, and Kaggle notebook are public, but JudgeBoi returned 502 on 2026-09-30 and the paper questions live on NTU COOL, so outside readers can only check accuracy inside the notebook.
HW8 involves no coding and no code submission. The TAs provide a finished Colab that runs Llama-3.2-1B-Instruct on the first 100 GSM8K questions and compares direct inference, Self-Consistency, Self-Certainty, and DeepConf (Confidence), sampling 16 reasoning traces per method. You read three papers, run the notebook, and answer 20 questions on NTU COOL: 18 about the papers and 2 about the Colab results. The prerequisite is Lecture 7 (Reasoning) of Lee's 2025 course. All questions are printed in hw8.pdf in Chinese and English, and the Colab is publicly downloadable. Only the COOL quiz and grades need an NTU account.
HW9 has 19 questions worth 10 points, answered only on NTU COOL with no code submission. The first 16 cover four papers — DDPM, Flow Matching, Rectified Flow, and MeanFlow — ending with questions that compare their training signals and few-step generation. The last 3 require the Colab: train two small MLPs on a 2D Swiss roll, one Flow Matching model that learns instantaneous velocity (always evaluated with 50 Euler steps, converged at Histogram JS ≤ 0.10) and one MeanFlow model that learns average velocity (always one-step, ≤ 0.40). Then compare 1 step vs 1 step, Flow Matching across Euler step counts, and Euler vs RK4 at equal steps and at similar compute. The PDF includes a generative-modeling tutorial that skips most of the math, and every question is published in Chinese and English. Outside readers miss only the COOL grading and answers.
KV Cache stores the keys and values already computed so decode does not recompute them, but every token costs memory. For Gemma 2 27B that is about 0.72MB per token, so an 80GB A100 holds only about 114k tokens. Hung-yi Lee then walks through ways to shrink it: let queries share keys and values (MQA, GQA), compress keys and values into one vector without ever decompressing (MLA), limit the attention span (Sliding Window, StreamingLLM), and drop keys and values nobody attends to (Scissorhands, H2O). He ends with cross-conversation prompt caching: it only hits when the prefix is identical, so a system prompt should put stable content first.
Hung-yi Lee opens his May 8 lecture by admitting that "self-improving AI" has no clear definition: it is a process of humans gradually letting go. He splits machine learning into three steps and checks where the "I" can be replaced by AI. Answers can come from the model's own self-corrections, reward shaping can be written by an LLM, the loss can be set by the model itself (scores, majority vote, entropy), and even the questions can come from a proposer model. But experiments keep showing that with no human at all, progress plateaus or the model trains itself into the ground. A strong AI can already train a weaker one, just not better than humans do. His verdict: in May 2026, AI is "still standing at the bank of the Rubicon."
Part 1 was about an AI setting its own loss and updating its own parameters. Part 2 fills in the other half: AI Agent = Harness + LLM, and the harness can grow too. You can't take a gradient through a harness, so the usual move is to hand it to a language model as a rewriter and keep a pool of candidates, much like a genetic algorithm (OPRO, GEPA, Darwin Gödel Machine; DSPy if you want a ready-made tool). Three extensions follow: updating harness and parameters together beats updating either alone; when the goal changes you have to choose between discarding everything and carrying everything, and editing a harness can cause forgetting too; and the update rule itself can be updated (HyperAgent, Gödel Agent, SEAL), which is meta learning. Hung-yi Lee closes with a new analogy — parameters are genes, context is the neurons — then argues that today's agents lack intrinsic motivation, and that the likeliest source of runaway growth is a gap between the goal humans meant and the goal the AI inferred.
National Taiwan University spreads its AI/ML courses across Electrical Engineering and Computer Science, and an official 'Machine Learning and AI' specialization stacks them into four levels. Hung-yi Lee's courses are the most open: ML 2026 Spring and Intro to GenAI and ML 2025 Fall publish slides, recordings, homework PDFs, and Colab notebooks, with only grading held back. Hsuan-Tien Lin's lectures are fully recorded and his Fall 2024 HW0–HW7 remain on the course page; Yun-Nung Chen's lectures are fully recorded, but most homework is public only as walkthrough videos; and since August 2025 Coursera only lets free learners watch the first module.