Skip to content

Hsuan-Tien Lin's ML Techniques T16 Finale: Three Families of Techniques, Plus Fall 2024's Modern Deep Learning Slides

Sep 30, 20261 min
TL;DRTechniques T16 re-sorts the whole course into three families: how to exploit features (kernels, aggregation, extraction, low-dimensional compression), how to optimize (gradients, equivalent problems, multiple steps), and how to fight overfitting (regularization, validation). It then uses four KDD Cup–winning models to show how the pieces combine in practice. The MOOC was recorded in 2016 and its deep learning stops at pre-training. The Fall 2024 on-campus course filled the gap with 302u (the ReLU family, Xavier/He initialization), 303u (momentum, RMSProp, Adam), a 2020 keynote deck, mlmai.ics, and 1126, an 11-model summary. The Fall 2026 versions of these files are scheduled for week 16 and currently return 404.

🌏 中文版

This is part 16 of the Reading NTU Hsuan-Tien Lin Machine Learning Foundations & Techniques series, following RBF Networks, k-Means, and Matrix Factorization. It has two parts: Lecture 16, Finale, of Machine Learning Techniques, and the four decks the Fall 2024 on-campus course added in its last three weeks.

Official material used:

Access level: all four Fall 2024 decks download freely, but they are slides only, with no recordings. Fall 2024 HW7 was released in week 14 and goes up to the neural networks of T12; it does not test 302u or 303u. So this stretch is A2. The access levels are defined in the Global AI/CS course map.

Three versions, three different endings

The three versions of the course end differently, so start by laying them out:

VersionFinal weeks
MOOC (2016)T13 deep learning (pre-training, autoencoders) → T14 RBF → T15 matrix factorization → T16 Finale
Fall 2024W13 (11/25) 212u, 213u → W14 (12/02) modern deep learning: 302u, 303u → W15 (12/09) no class, mlmai.ics posted → W16 (12/16) finale: 1126
Fall 2026W13 (12/02) no class, mlmai.ics → W14 (12/09) final exam → W15 (12/16) 212u, 213u → W16 (12/23) modern deep learning and finale: 302u, 303u, 1126

Two things stand out. First, neither on-campus semester uses the MOOC's 216 slides; finale week uses 1126 instead. Second, the Fall 2026 paths for 302u, 303u, 1126, and mlmai.ics all returned 404 on 2026-09-30. They go up in their scheduled weeks. This post therefore follows the Fall 2024 versions.

MOOC T16: three families of techniques

T16 teaches nothing new. It re-sorts the first 15 lectures along a different axis. The final summary slide has three lines: kernel, aggregation, extraction, low-dimensional; gradient, equivalence, stages; (lots of) regularization, validation.

Family 1: exploiting features

Slides 2–5 use four tables to answer "where does the feature transform Φ live?"

  • Kernel: many features are embedded in the kernel's inner product. The polynomial kernel is a scaled polynomial transform, the Gaussian kernel an infinite-dimensional one, and the stump kernel uses decision stumps as the transform. Adding kernels takes the union of transforms; multiplying them combines transforms. Models that use them: SVM, SVR, kernel ridge regression, kernel logistic regression, probabilistic SVM.
  • Aggregation: each gt is itself a predictive feature φt(x). The base unit can be a decision stump, a decision tree, or a Gaussian RBF. The combination is uniform (bagging, random forest), non-uniform (AdaBoost, gradient boost), or conditional (decision tree, nearest neighbor).
  • Extraction: features are hidden variables learned jointly with the usual weights, often with help from unsupervised learning. NNs extract neuron weights, RBF networks extract centers, and matrix factorization extracts user and movie factors; k-means supplies cluster centers, and autoencoders and PCA supply basis directions.
  • Low-dimensional compression: compress the original features into fewer. Autoencoders and PCA compress while preserving information. Decision stumps and tree branches are the "best" naive projection to one dimension. Random forest branching is a "random" low-dimensional projection. Matrix factorization projects abstract features into concrete ones. Feature selection is the "most helpful" low-dimensional projection.

Family 2: optimizing

Slides 7–9 name three strategies:

  1. Gradient descent: whenever ∇E is roughly defined, use the first-order update "new variables = old variables − η∇E." SGD, minibatch, and full GD handle kernel logistic regression, NNs (backprop), and matrix factorization. Steepest descent and functional GD handle AdaBoost and gradient boost.
  2. Solve an equivalent problem: when the original is hard, find an equivalent one. The dual SVM relies on convex QP, kernel logistic and kernel ridge regression rely on the representer theorem, and PCA is equivalent to an eigenproblem.
  3. Split into steps: break the hard problem into easier sub-problems. Multi-stage: probabilistic SVM, linear blending, stacking, RBF networks, DeepNet pre-training. Alternating optimization: k-means, alternating least squares. Divide and conquer: decision trees.

The Fun Time on slide 10 ties all three together. Running T13's DeepNet on PCA-preprocessed data uses every strategy: minibatch GD for training, an equivalent eigenproblem for PCA, and multiple stages for pre-training.

Family 3: eliminating overfitting

Slides 11–12 split this into two kinds:

  • Regularization, which the slides call "arguably the most important technique": large margin (SVM, and AdaBoost indirectly), L2 (SVR, kernel models, weight decay in NNs), voting and averaging (uniform blending, bagging, random forest), denoising (autoencoders), weight elimination and early stopping (NNs), constraining (tied autoencoder weights, the number of RBF centers), and pruning (decision trees).
  • Validation, "simple but necessary": the number of support vectors for SVM/SVR, the OOB error for random forests, and internal validation for blending and decision tree pruning.

Section 4: combining them in practice

Slides 14–17 walk through four KDD Cup titles won by NTU teams:

CompetitionModel highlights (as summarized on the slides)
KDD Cup 2010 (Yu et al.)Logistic regression on many raw-encoded features plus random forest on human-designed features, linearly blended
KDD Cup 2011 Track 1 (Chen et al.)Matrix factorization variants, RBMs, k-NN, PLSA, linear regression, NNs, GBDT, combined with NNs, decision-tree-like models, and linear blending
KDD Cup 2012 Track 2 (Wu et al.)Variants of linear regression, logistic regression, and matrix factorization, combined with NNs, GBDT-like models, and linear blending; the key is to blend without overfitting
KDD Cup 2013 Track 1 (Li et al.)Random forests with very many trees and GBDT variants, plus a large effort on features built from domain knowledge

For the 2010 entry the slide says, "yes, you've learned everything!" Slide 18 lists the ICDM 2006 top 10 data mining algorithms (C4.5, k-means, SVM, Apriori, EM, PageRank, AdaBoost, k-NN, Naive Bayes, C&RT), then adds five ML methods Lin personally thinks are missing: linear regression, logistic regression, random forest, GBDT, and NNs. Slide 19 is a word cloud titled "welcome to the jungle."

Try this: on a sheet of paper, make the three families the columns and write in every model you remember. The empty cells tell you which lectures to rewatch. This table checks whether you connected the course better than any set of notes.

Fall 2024 supplement 1: 302u, activation and initialization

The 302u title slide reads "Machine Learning Soundings (機器學習深測), Lecture 2: Activation and Initialization." The course page labels it "deep learning activation." It follows 212u/213u and starts from backprop in a tanh network.

Vanishing gradients (slides 3–6): in backprop, δ(1) is δ(L) multiplied by a long chain of weights and ϕ′(s) terms. tanh saturates when |s| is large, so ϕ′(s) = 1 − tanh²(s) goes to 0. The early layers get gradients too small to update, and deep networks stop training. Slide 6 lists six remedies: skip connections, small random initialization, layer-wise pre-training (Techniques T13), internal normalization, gradient normalization, and better activation functions.

The ReLU family (slides 7–10):

  • ReLU, ϕ(s) = max(s, 0): the derivative is 1 on the positive side, so gradients avoid vanishing about half the time. Each example activates only part of the network, and the arithmetic is fast. The slides call it arguably the most widely used activation in deep learning.
  • Dead neurons: if a neuron's s is negative for every example, its output and gradient are both 0 and it never updates again. One very large gradient step can push the bias far negative; so can all-positive inputs (unshifted images) paired with negative weights.
  • Leaky ReLU, ϕ(s) = max(s, 0.01s): a small slope on the negative side makes dead neurons less likely. The slide itself asks: why 0.01?
  • Parametric ReLU, ϕ(α, s) = max(s, α·s): learn α with backprop too. The slide's takeaway: "anything (loosely) differentiable is learnable."

Initialization (slides 11–14): all zeros is too symmetric for tanh and not differentiable for ReLU; a shared constant clones the neurons; values that are too large saturate tanh and kill or overfit ReLU units. So you want small, zero-mean random values. The slides then derive the variance:

  • For a tanh forward pass that keeps each layer's output variance constant, var(w) = 1/d(ℓ−1); the backward pass wants 1/d(ℓ). Xavier initialization compromises with var(w) = 2/(d(ℓ−1) + d(ℓ)).
  • ReLU keeps only half, so E(x²) is half the variance of s, and you multiply back by 2: He initialization, var(w) = 2/d(ℓ−1).

Original papers listed on the course page: ReLU, Glorot, Bordes & Bengio; Leaky ReLU, Maas, Hannun & Ng; Parametric ReLU, He, Zhang, Ren & Sun.

Fall 2024 supplement 2: 303u, optimization in deep learning

303u is Lecture 3 of the same series, Optimization in Deep Learning, and it is short.

Slide 2 describes the error surface: local minima are not as bad as imagined; saddle points and local maxima are easy to escape, especially with SGD; plateaus need a larger η; ravines need protection against oscillation. Backprop makes gradients slow to compute, so you use minibatch SGD, and the price is a noisy gradient estimate. To stabilize it, average.

  • Momentum (slides 3–5): averaging M minibatches costs M times the computation, but reusing past gradients does not. A uniformly weighted moving window is not ideal, so use an exponentially decaying average, vt = βvt−1 + (1 − β)Δt; β = 0 recovers plain SGD. It cancels some variance, damps oscillation across ravines, and escapes shallow local optima and saddle points.
  • RMSProp (slides 6–8): take smaller steps along components with larger gradients. With only stochastic gradients available, keep a moving average ut of Δt ⊙ Δt, and set each component's step to η · (ut + ϵ)−1/2.
  • Adam (slide 9): roughly momentum plus RMSProp plus a global decay. The slide warns that Adam is usually more aggressive than plain SGD, but it can also overfit faster.

Papers listed on the course page: backprop by Rumelhart, Hinton & Williams, momentum by Qian, and Adam by Kingma & Ba.

Try this: write a two-layer ReLU network in numpy. On one dataset, compare three initializations (all zeros, standard normal, He) against three optimizers (SGD, momentum, Adam), and plot one training-loss curve per combination. Nine curves on one chart is the full experimental version of 302u plus 303u.

Fall 2024 supplement 3: the mlmai.ics keynote

In week 15 of Fall 2024, class was cancelled because the instructor was attending ACML 2024 and NeurIPS 2024, and the course page posted mlmai.ics instead. It is not a lecture handout. It is Lin's keynote "Machine Learning for Modern Artificial Intelligence," given on 2020-12-17 at the International Computer Symposium 2020, held jointly with a Taiwan Ministry of Education AI talent program showcase.

The talk has three parts:

  1. ML for (Modern) AI: Lin defines modern AI as "intelligently ≈ easily," or application intelligence. Machine learning is the tool that "cooks" big data into AI, and early on it often has to lean on human domain knowledge. The example is estimating tropical cyclone intensity with a CNN.
  2. ML Research for Modern AI: three research cases. Cost-sensitive multiclass classification (mistakes have different costs, such as predicting COVID-19 as healthy); active learning by learning, which uses a bandit to pick among active learning strategies and was released as libact; and cyclone intensity estimation. Each case ends with a "have we made it more realistic?" reflection. For libact, the most-reported issue was that it is hard to install on Windows and Mac.
  3. ML for Future AI: more creative, more explainable, more interactive, mapped to winning human respect, trust, and heart.

The deck is worth reading as a demonstration of how textbook ideas, like cost-sensitive error (Foundations L8) and bandits (the slot machine problem in Foundations HW2), turn into research questions.

Fall 2024 supplement 4: the 1126 super short summary

The 1126 title slide reads "Machine Learning Foundations, Lecture 1126: Super Short Summary." (1126 is "the official lucky number of this class," per the last Fun Time in T16.) It condenses the semester into 11 models, each described with the same fields: err, the error the algorithm actually optimizes, optimization method, Φ, regularization/validation, parameters, and practical use.

#ModelPractical use, per the slides
1PLAonline learning, and "teaching"
2ridge linear regressiona decent baseline
3logistic regressionbaseline for hard and soft classification
4ridge polynomial regressionoften for 1-D regression
5soft-margin linear SVMlarge-scale classification
6soft-margin kernel SVMmid-sized classification
7AdaBoost"boosting" decision trees or stumps
8decision treean explainable nonlinear model
9bagging/random foreststabilizing any model or tree
10GBDTinformation retrieval and competitions
11neural networks/deep learningvision and speech

The entry for model 11 pulls in 302u and 303u directly: GD/SGD with Adam, Xavier/He initialization, ReLU or tanh neurons, and regularization by early stopping, L1/L2, and dropout.

Try this: read this table next to the T16 taxonomy. 1126 is organized by model and T16 by technique; cross the two and you have an index to the whole course.

Videos

T16 Finale (MOOC; the four Fall 2024 supplementary decks have no matching recordings):

The Fall 2026 course page advertises a public live stream, but I did not check whether complete recordings stay up afterward.

Next

The course content ends here. The next post, Foundations Homework Guide: Fall 2024 HW0–HW5, goes back over the Foundations homework, followed by Techniques Homework and the Final Project.

Further reading: 302u and 303u are only a doorway into deep learning. To go further, this site has CMU 11-785 (its lectures on loss surfaces and momentum and optimizers and regularization pick up right where 303u stops), MIT 6.7960, and NTU's other track, Hung-yi Lee's ML 2026.

References