The first self-study deck of ADL Fall 2025 describes machine learning as finding a function from data, and deep learning as a production line of simple functions where the machine learns what every station does. It uses speech and vision to contrast deep and shallow models, credits big data and GPUs for the post-2010 breakthroughs, and uses the universality theorem to ask why networks should be deep rather than fat. The most practical part comes last: the output domain decides the learning task, and the architecture should fit the properties of the input domain.
Lectures 7 and 8 of Machine Learning Techniques open the aggregation part of the course. T7 sorts ways of combining hypotheses into uniform, linear, and any blending (stacking), shows with a few lines of algebra that uniform blending reduces variance, and then uses the bootstrap to create diverse g_t from the single dataset you have: that is bagging. T8 reinterprets the bootstrap as example weighting, then deliberately up-weights the examples the previous hypothesis got wrong so the next one is forced to differ, and votes with α_t = ln √((1−ε_t)/ε_t): that is AdaBoost. Practice with Fall 2024 HW6 Q4 and Q9, plus HW7's bootstrap and AdaBoost proofs and a 500-round AdaBoost-Stump experiment on madelon. There are no official solutions.
Hsuan-Tien Lin's Machine Learning Foundations (16 lectures) and Machine Learning Techniques (16 lectures) are two Mandarin-taught MOOCs. All 130 YouTube videos and 32 slide decks are free. The MOOCs alone are A2: since August 2025, free Coursera accounts can only view the first module, so the exercises sit behind a paywall. Add the Fall 2024 course page, which publishes HW0–HW7 and the final project spec, and you reach A3, minus the grading chain: no official solutions, Gradescope and NTU COOL are enrolled-only, and the Kaggle competition returns 404. Fall 2026 is running now as a flipped classroom; slides through week 4, hw0, and hw1 are public.
Lectures 9–11 of Machine Learning Techniques tie three models together with one thread: trees plus aggregation. T9 treats a decision tree as conditional aggregation and covers C&RT's binary branching, Gini and regression impurity, pruning, categorical features, and surrogate branches. T10 applies bagging to fully grown trees; add random subspaces and random projections and you get a random forest, with free OOB validation and permutation-based feature importance. T11 re-derives AdaBoost as steepest descent in function space on the exponential error, then swaps in squared error to get GBDT, which fits regressions to residuals. Practice with the impurity and gradient boosting proofs in Fall 2024 HW7. There are no official solutions.
Foundations Lecture 4 first shows that learning is impossible: from D alone, any guess outside D can be called wrong. That is No Free Lunch. It then reframes the question with marbles in a bin. If the data is drawn independently from one distribution, Hoeffding's inequality says the in-sample error E_in is probably close to the true error E_out. Checking one fixed h is only verification. Once the algorithm chooses among M hypotheses, a union bound charges 2M exp(−2ε²N). Conclusion: with a finite hypothesis set and small E_in, learning is feasible. What to do when M is infinite is the next lecture's job.
Techniques T16 re-sorts the whole course into three families: how to exploit features (kernels, aggregation, extraction, low-dimensional compression), how to optimize (gradients, equivalent problems, multiple steps), and how to fight overfitting (regularization, validation). It then uses four KDD Cup–winning models to show how the pieces combine in practice. The MOOC was recorded in 2016 and its deep learning stops at pre-training. The Fall 2024 on-campus course filled the gap with 302u (the ReLU family, Xavier/He initialization), 303u (momentum, RMSProp, Adam), a 2020 keynote deck, mlmai.ics, and 1126, an 11-model summary. The Fall 2026 versions of these files are scheduled for week 16 and currently return 404.
Fall 2024 had six Foundations assignments. HW0 is 20 multiple-choice math prerequisite questions. HW1–HW5 each have 12 problems plus a bonus: Q1–4 are auto-graded, Q5–12 are graded by TAs, the programming problems use rcv1, cpusmall, and mnist from the LIBSVM datasets site, and HW5 uses LIBLINEAR. HW1 and HW2 each include a problem where you argue with a ChatGPT-style answer. Fall 2026 has released hw0 and hw1: hw1 is now 16 multiple-choice problems with 4 secretly chosen for TA grading, and the data is the course's own hw1_train.dat. The new policy allows AI tools and vibe coding, but AI-generated code needs block-by-block comments in your own words. Neither semester publishes official solutions.
Lecture 5 of Machine Learning Techniques rewrites the soft-margin SVM in unconstrained form: ½wᵀw plus C times the total hinge error. That is an L2-regularized model, and a larger C means weaker regularization. The hinge error and logistic regression's cross-entropy are both convex upper bounds of the 0/1 error, so the SVM approximates L2-regularized logistic regression. For probability outputs, you can use Platt's two-level learning, running logistic regression on top of SVM scores, or use the representer theorem to do kernel logistic regression directly. Lecture 6 uses the same theorem to get the closed form β = (λI + K)⁻¹y for kernel ridge regression, but β is dense; switching to the ε-insensitive tube error gives SVR with sparse coefficients. Fall 2026 does not schedule these two lectures.
Lecture 3 of Machine Learning Techniques merges "feature transform + inner product" into a single kernel function K(x, x′). Training and prediction in the dual SVM only need K, so d̃ can be infinite: the Gaussian kernel corresponds to an infinite-dimensional transform. Lecture 4 admits the SVM can still overfit and introduces violations ξₙ and a parameter C, giving the soft-margin SVM. Its dual differs from the hard-margin one in exactly one way: αₙ gets an upper bound C. The value of αₙ sorts the data into non-SVs, free SVs, and bounded SVs, and the fraction #SV/N upper-bounds the leave-one-out error, a cheap way to rule out dangerous (C, γ).
The first three lectures of Machine Learning Foundations define machine learning as a flow chart: an unknown target function f generates data D, and an algorithm A picks g from a hypothesis set H, hoping g ≈ f. The simplest H (the perceptron) and A (PLA) then show the chart in action. On linearly separable data, PLA makes at most R²/ρ² updates; on non-separable data, use pocket instead. Lecture 3 sorts learning problems along four axes: output, label, protocol, and input. Foundations mostly deals with batch, supervised binary classification or regression on concrete features. Practice with Fall 2024 HW1 and Fall 2026 hw1.
Lecture 11 of ML Foundations compares PLA, linear regression, and logistic regression on the same score s = wᵀx. The three differ only in their error functions, and scaled cross-entropy upper-bounds the 0/1 error, so both regressions can do classification. The lecture then turns logistic regression into SGD by computing the gradient on one random example, and builds multiclass classifiers from binary ones with OVA and OVO. Lecture 12 uses a feature transform Φ to turn a circular boundary into a line in Z-space. The price is that computation and d_vc both grow with the dimension, so the advice is: try a linear model first. Practice problems are in Fall 2024 HW4.
Lecture 1 of Machine Learning Techniques turns "which separating line is best?" into an optimization problem. Once you fix the scale so that min yₙ(wᵀxₙ+b) = 1, maximizing the margin is the same as minimizing ½wᵀw, which is a standard QP. Lecture 2 uses Lagrange duality to trade a QP with d̃+1 variables for one with N variables and N+1 constraints, then uses the KKT conditions to recover (b, w) from α. Only the points with αₙ > 0, the support vectors, affect the answer. The dual still contains the inner product zₙᵀzₘ, so the dependence on dimension is not really gone until the kernel lecture.
Linear regression writes squared error as (1/N)‖Xw − y‖², sets the gradient to zero, and gets w_LIN = X†y in one step. The hat matrix H = XX† projects y onto the column space of X, which shows that on average E_out − E_in ≈ 2(d+1)/N. Logistic regression estimates P(+1|x) with θ(wᵀx); maximum likelihood turns into the cross-entropy error ln(1 + exp(−y wᵀx)). It has no closed-form solution, so you walk downhill along −∇E_in step by step. That is gradient descent.
Lectures 12 and 13 of Machine Learning Techniques open the third part, distilling hidden features. T12 starts from a linear combination of perceptrons: two layers can build AND and OR but not XOR, and one more layer fixes that, which is the multi-layer perceptron. It then replaces sign with tanh, derives backprop, and covers non-convex optimization, d_vc = O(VD), weight elimination, and early stopping. T13 discusses the challenges of deep networks, uses autoencoders as information-preserving encodings for layer-wise pre-training, treats denoising as regularization, and proves that the optimal linear autoencoder is spanned by the top eigenvectors of XᵀX, which is PCA. The videos date from 2016; modern deep learning is covered by the Fall 2024 302u/303u slides. Practice: Fall 2024 HW7 Q4, Q9, and bonus Q13.
Lecture 13 of ML Foundations defines overfitting as 'lower E_in but higher E_out' and uses experiments to find four causes: too little data, stochastic noise, an overly complex target (deterministic noise), and excessive model power. Lecture 14's remedy is regularization. It rewrites 'step back to H₂' as the constraint ‖w‖² ≤ C, then uses a Lagrange multiplier to turn it into minimizing E_in + (λ/N)wᵀw, which is weight decay. Back in VC theory, regularization shrinks the effective VC dimension d_EFF, and L1 buys sparse solutions. Practice problems: Fall 2024 HW4 Q8–9 and HW5 Q1, Q5–6, Q10.
Techniques T14 reinterprets the Gaussian SVM as a linear vote over distance-based similarities, which gives the RBF network. Too many centers overfit, so k-means picks a few prototypes, and k-means itself is alternating optimization. T15 starts from the Netflix ratings data: one-hot encode user IDs, feed them into a linear network with the tanh removed, and you get matrix factorization R ≈ VᵀW, learned by alternating least squares or SGD. The lecture closes with a map of extraction models: boosting, neural nets, RBF networks, matrix factorization, and k-NN. These two lectures exist only as MOOC material. Neither the Fall 2024 nor the Fall 2026 schedule covers them, and no public homework problem does either.
The Techniques half of Fall 2024 has two homework sets and a final project, and all three PDFs are public. HW6 covers kernels, soft-margin SVM, and aggregation; its programming part uses LIBSVM on the 3-vs-7 subproblem of mnist.scale to count support vectors, compute margins, and run 128 validation rounds. HW7 covers bootstrap, impurity, AdaBoost, gradient boosting, and neural networks; its programming part is a 500-round AdaBoost-Stump on madelon. The final project is a fictional baseball league, HTMLB: predict home-team wins across two Kaggle stages and write an English report of at most seven pages that compares at least four methods. There are no official solutions. On 2026-09-30 both Kaggle pages returned 404 without login, so outside readers probably cannot get the HTMLB data and should reproduce the same splits on a public dataset instead.
The Hoeffding guarantee from L4 carries an M, the number of hypotheses. Perceptrons have infinitely many lines, so M blows up. L5 stops counting hypotheses and counts how many ○× patterns (dichotomies) they can produce on N data points instead; the maximum is the growth function m_H(N). 2D perceptrons produce at most 14 patterns on 4 points, fewer than 2⁴ = 16, so 4 is their break point. L6, marked optional by the course, proves that any break point caps m_H(N) by a polynomial, which is what makes the VC bound work.
Lecture 15 of ML Foundations tackles model selection. Selecting by E_in overfits, and selecting by E_test is cheating. The compromise is to carve a validation set out of the training data, select by E_val, then retrain on all the data. The validation size K is a dilemma, with K = N/5 as the rule of thumb. Leave-one-out is almost unbiased but expensive and unstable, so in practice you use 5-fold or 10-fold. Lecture 16 closes with three principles, Occam's razor, sampling bias, and data snooping, and a 'Power of Three' recap: three related fields, three bounds, three linear models, three tools. Practice problems are in Fall 2024 HW5.
L7 names the largest non-break point the VC dimension d_VC, proves that d-dimensional perceptrons have d_VC = d + 1, and rewrites the VC bound as E_out ≤ E_in + a model-complexity penalty, so both too large and too small a d_VC hurt. Theory asks for N ≈ 10,000·d_VC examples; in practice 10·d_VC is often enough. L8 swaps the fixed target function for a distribution P(y|x) and shows the VC theory still holds under noise. The error measure should come from the application: a CIA fingerprint check that penalizes admitting an intruder 1000 times more can be reduced to plain classification by copying examples.
Hung-yi Lee's Spring 2026 Machine Learning course at National Taiwan University opens with OpenClaw. The first half takes apart AI agents, context engineering, inference speed-ups, and positional embeddings. The second half covers harness engineering, self-correction, and self-improving AI. Slides and recordings for all 8 lectures, plus PDFs and Colab notebooks for all 10 assignments, are public, so it rates A3. What's missing is grading: JudgeBoi returned 502 on 2026-09-30, NTU COOL is campus-only, and the three guest talks have no materials at all.
National Taiwan University spreads its AI/ML courses across Electrical Engineering and Computer Science, and an official 'Machine Learning and AI' specialization stacks them into four levels. Hung-yi Lee's courses are the most open: ML 2026 Spring and Intro to GenAI and ML 2025 Fall publish slides, recordings, homework PDFs, and Colab notebooks, with only grading held back. Hsuan-Tien Lin's lectures are fully recorded and his Fall 2024 HW0–HW7 remain on the course page; Yun-Nung Chen's lectures are fully recorded, but most homework is public only as walkthrough videos; and since August 2025 Coursera only lets free learners watch the first module.
Several semesters of Berkeley CS189 are online at once. From here on, this series follows Spring 2026 (Listgarten/Dimakis): its slides, 25 lecture videos, discussions with solutions, HW1–5 handouts and midterm solutions are all publicly accessible, so it rates A3. Spring 2025 (Shewchuk) is a different, classic route and the only one that covers SVMs, decision trees, PCA and boosting. Fall 2025's homework folders open empty to outside readers, so it was not chosen. Fall 2026 is still running and serves only as a comparison.
CS189 Spring 2026 HW1 has three pieces: ten written math warm-ups (a linear system, limits of matrix powers via eigendecomposition, SVD, a matrix that flips an image, partial derivatives, a chain rule over a recursion, and four probability problems including Bayes for cancer screening), plus two public Modal notebooks on Fashion-MNIST. Part 1 drills pandas / Plotly / K-means / an MLP / matrix-based image augmentation / tensor puzzles; Part 2 does price regression, MAE / MSE / R², confusion matrices, and finally a secret test set whose images have been rotated. Due Feb 20; no official solutions.
HW2 is a written-only assignment with 10 problems. The first half practices reading a paper (Chatbot Arena) and the core derivations for regression and MLE/MAP. The two heaviest problems come last: Mixed Feelings goes from k-means' fragility to outliers through robust k-means to a GMM with a uniform background, and Watch Me Flow Dat proves that conditional flow matching and a discretized MLE are the same objective. The problem PDF and LaTeX template are freely downloadable; there are no official solutions.
The first three lectures of CS189 Spring 2026 hold off on derivations. They teach you how to tell whether a problem calls for ML, how to look at data with pandas and Plotly, and how to run one full train/validate/test cycle in scikit-learn. Lecture 1 has slides but no recording; Lectures 2–3 have slides and video; Discussion 1 is a calculus, linear algebra and probability warm-up with solutions and a walkthrough. Together they set you up directly for HW1.
Lectures 4–7 of CS189 Spring 2026 tie unsupervised learning into one thread. K-means clusters the data, then its weaknesses show up: hard assignments, no probabilistic framing, and a bias toward round clusters of similar size. The course reviews probability, introduces maximum likelihood estimation (MLE) and multivariate Gaussians, and rewrites K-means as a Gaussian mixture model (GMM). The GMM log-likelihood has no closed-form solution, and that gap leads the course to gradient descent. All four lectures have slides and video; Discussions 2–3 come with solutions and walkthroughs.
Over four lectures, CS189 Spring 2026 presents linear regression from three angles that meet in one formula: MLE under Gaussian noise is least squares; the least-squares solution is the orthogonal projection of y onto the column space of X; and when features are collinear or too many, ridge (the MAP estimate under a Gaussian prior) or lasso (a Laplace prior) pulls the solution back, with λ chosen on a validation set. Slides, videos, and Discussion 3–4 solutions are all publicly accessible.
CS189 Spring 2026 Lec 11–12 splits classification into two routes. Generative models fit p(x|y) for each class (GDA: shared covariance gives LDA and a linear boundary, per-class covariance gives QDA and a quadratic one). Discriminative models fit p(y|x) directly (logistic regression: sigmoid, softmax, cross-entropy MLE, no closed form, so gradient descent). The bridge: the LDA posterior can always be written in logistic form, but not the other way around. For evaluation, accuracy misleads under class imbalance; ROC/AUC sweeps every threshold and ignores calibration; PR curves care about class balance.
CS189 Spring 2026 covers gradient descent in two lectures. Lec 13 derives the learning-rate limit and the condition number from the Hessian's eigenvalues, then moves through momentum, learning-rate schedules, AdaGrad/RMSProp/Adam, and mini-batch SGD. In Lec 15, Dimakis walks through the same material again, starting from a gradient computed by hand on a small data table. Both slide decks, the recordings, Lec 13's handwritten notes, and Discussions 6 and 7 (with solutions) are all publicly accessible.
Lec 14 recasts ridge as MAP under least squares plus a Gaussian prior, breaks down bias-variance, and closes by using Chatbot Arena to show how to read a paper. Lec 16 builds entropy from compression, moves on to KL and cross-entropy, and lands back on the logistic regression loss. The 3/17 midterm and its official solutions are public in the past-exams folder, with 6 problems, 56 points, 110 minutes, and 6 problem-by-problem walkthrough videos, so you can sit it as a mock exam.
07-380 Lec1 has no algorithms. It sets up three things: the working definition that intelligence means doing well on a task under uncertainty, a bubble diagram that color-codes 07-280 topics against the new 07-380 topics, and a grading scheme with quizzes at 55%, no final exam, and a final project. Off-campus readers can use it to place the other 25 lectures.
Lecture 8 swaps MLE's argmax p(D|θ) for argmax p(θ|D). The prior p(θ) multiplies the likelihood, and after taking the negative log it becomes an extra term in the objective. A trick coin shows data overwhelming the prior, a Beta prior estimates a click rate, and Gaussian and Laplace priors on linear-regression weights turn into L2 and L1 regularization.
Lecture 9 stops learning p(y|x) directly. Instead it learns the class prior p(y) and the class-conditional p(x|y), then inverts them with Bayes rule. The price is stronger assumptions; the payoff is the ability to generate new data and more stability with little data. Naive Bayes uses conditional independence to make the parameters estimable, GDA uses multivariate Gaussians for continuous features, and whether the covariances match decides a linear or a curved boundary.
The public materials for the CS181 2026 final (May 9) — the final checklist, 16 second-half practice problems, and a 66-page final review — are all 2025 versions. They cover Bayes nets and EM, which the 2026 schedule never lists, and skip Transformers, VAEs, GANs, and autoregressive models. Sort the checklist against the 2026 schedule into three buckets, cover the new topics with homework and sections, then do the 2025 practical for one end-to-end project.
HW4 Problem 2 has you train a convolutional autoencoder on 64×64 CelebA, sample from N(0, I), and watch it fail to produce faces. You then derive the ELBO, the reparameterization trick, and the closed-form KL, and turn the same backbone into a VAE to compare reconstructions and samples.
HW5 Problems 3–4 give handwritten digits to three methods that never see a label: K-means summarizes the data with 10 mean images, HAC builds a merge tree you can cut at any number of clusters, and PCA compresses images onto a few continuous directions. All three answer the same question — how much error do you pay to describe the data with a few objects — and the assignment makes you compare their objectives and reconstruction errors directly.
HW5 Problems 1–2 both turn learning into a classification task. SimCLR's NT-Xent loss asks the network to pick the other augmented view of the same image out of 2N−1 candidates; a GAN's discriminator classifies real versus fake. You derive the math behind each (the cross-entropy equivalence, the optimal discriminator and the JS divergence), then write both training loops on FashionMNIST and MNIST.
HW6 Problem 4 (20 points) takes apart the cost of generating one token at a time in three questions: picking the most likely token at each step doesn't give the most likely sequence; recomputing every key at every step makes cost quadratic, and a KV cache brings it back to linear; speculative decoding lets a small model guess and a large model verify in one pass. It is all pencil-and-paper, and every question maps onto a real design choice in today's LLM inference systems.
The CS181 Spring 2026 midterm is in class on Mar 10, worth 15% of the grade, closed-book with one double-sided note sheet. The official midterm checklist has four blocks: regression, classification, neural networks and model selection, and SVMs. This post maps each block to HW0–HW3 problem numbers, flags what the 2026 homework never drilled, and explains how to use the 2025 practice exam, review session, and concept checks.
Today's ML Fundamentals rotation covers five core concepts: what bagging and boosting each reduce (variance vs. bias), why boosting overfits with too many rounds while random forests don't, why XGBoost's second-order approximation (gradient plus Hessian) and objective-function-level regularization swept tabular ML, when to normalize vs. standardize a feature (it depends on the algorithm's assumptions, not the data), and why AdamW decouples weight decay from the L2 penalty term in the loss. The practice question comes from a real C3 AI Data Scientist interview: name the classic bagging algorithm, explain why bootstrap aggregation reduces variance, then answer a follow-up on what boosting trades off instead — the breakdown threads all five concepts into one ensemble-method decision framework.
Today's ML Fundamentals rotation covers five core concepts: diagnosing the bias-variance trade-off with learning curves rather than reciting definitions, how L1 and L2 regularization each add a different penalty term to the loss function, the trade-off between batch and stochastic gradient descent, why accuracy is the most misleading metric on imbalanced data (look at precision/recall/F1/ROC-AUC instead), and the data-leakage trap hiding inside cross-validation. The practice question is adapted from a real Transunion Data Scientist interview: design the feature engineering, model choice, and validation strategy for an end-to-end ML project on large-scale data — the breakdown threads all five core concepts into one decision framework.
Today's ML Fundamentals rotation covers the bias-variance tradeoff, the difference between L1 (Lasso), L2 (Ridge), and Elastic Net regularization, how to detect and fix overfitting versus underfitting, when to reach for K-fold, stratified, or time-series cross-validation, and which of precision, recall, F1, or AUC-ROC to prioritize under different business costs. The practice question is a classic: a model performs well offline but regresses after deployment — interviewers want to see a systematic debugging framework, not a guess.
This round of ML Fundamentals skips the textbook formulas and drills the diagnostics that separate senior candidates from junior ones: reading the train/val gap on a learning curve to decide whether to add capacity or regularize, the three usual suspects when CV and production scores diverge (StratifiedKFold, GroupKFold, TimeSeriesSplit), the cost-ordered playbook for class imbalance — reweight before you reach for SMOTE — and why picking a decision threshold from FP/FN cost is a separate problem from probability calibration.
ML fundamentals interviews test whether you can diagnose the gap between 'the metric looks great' and 'production is on fire.' Today covers four high-frequency topics: why AUC-ROC inflates under heavy class imbalance, why cross-entropy beats MSE for classification (it comes down to vanishing gradients), whether bagging or boosting fixes variance versus bias, and the common misconception that multicollinearity hurts prediction — it only hurts interpretability.
Week 4 enters ML: supervised classification (k-NN, SVM, Perceptron), model evaluation, RL basics (MDP, Q-learning, ε-greedy). Projects: Shopping (purchase prediction with k-NN) and Nim (learning to play via Q-learning).
A/B testing turns a product change into an estimate with uncertainty. A useful report covers effect size, confidence, guardrails, randomization, and launch risk.
Large-sample normal approximation describes the behavior of estimators, not raw data. It is useful, but dependence, boundaries, and distribution shift can make it unreliable.
Bayesian inference updates uncertainty about an unknown parameter by combining prior belief with the likelihood from observed data, producing a posterior distribution.
Bias checks whether an estimator is centered correctly, variance checks sampling fluctuation, MSE combines both, and consistency asks whether the estimator approaches truth as sample size grows.
Bootstrap estimates uncertainty by resampling from the observed sample with replacement, rebuilding many sample-like datasets, and watching the statistic fluctuate.
Chi-square tests compare observed counts with expected counts. First decide whether the problem is goodness-of-fit for one categorical variable or independence for two categorical variables.
Distributions are names for data-generating situations, not formula cards. Learn when Bernoulli, Binomial, Poisson, and Normal distributions fit a problem.
A confidence interval puts a point estimate back inside sampling fluctuation. Computing bounds is only the first step; you also need to explain standard error, critical values, and coverage.
Data type determines the statistical tools you can use. Start with categorical, numeric, count, and time-ordered data, then choose summaries that fit the question.
At the final review stage, train problem recognition: identify data type, unknown quantity, and decision goal before choosing a formula and writing a contextual conclusion.
Expectation describes long-run center; variance describes fluctuation. This post computes E[X], E[X^2], and Var(X), then connects them to average loss and model stability.
Experimental design decides whether a result can be interpreted. Randomization, control, blocking, replication, blinding, and pre-specified outcomes give inference a usable foundation.
Fisher information uses likelihood curvature to measure how well the data locate a parameter; larger information usually means a smaller standard error for the MLE.
A confidence interval is built by defining the target estimate, describing its sampling error, and choosing a rule that turns uncertainty into a range.
A generalized linear model starts from the response type, chooses a suitable distribution, and uses a link function to connect the mean to a linear predictor.
A hypothesis test is a decision process under uncertainty: write H0/H1, choose alpha, compute a test statistic and p-value, then decide whether the data is strong enough to challenge H0.
The inference map starts with the question type: point estimate, uncertainty interval, decision test, likelihood model comparison, Bayesian update, or resampling.
The likelihood-ratio test compares the log likelihood of a restricted model with a full model; the usual chi-square reference only makes sense under nested-model and approximation conditions.
OLS is a useful baseline, but coefficient interpretation, inference, prediction, and diagnosis depend on assumptions about linearity, errors, independence, and variance.
Logistic regression estimates probabilities first. Classification decisions come later, when thresholds turn those probabilities into actions under real error costs.
Logistic regression connects a linear score to a probability between 0 and 1. Understanding odds, log odds, and odds ratios prevents wrong coefficient interpretations.
MAP maximizes the posterior. After taking logs, the prior becomes a penalty term, which connects Bayesian estimation to L1, L2, and regularized ML objectives.
MLE fixes the observed data and compares which parameter values make that data most plausible; log likelihood turns products into sums and connects directly to negative log loss.
Method of Moments matches sample moments to theoretical population moments, then solves for parameters. It is not always the most efficient method, but it builds the first intuition for parameter estimation.
Nonparametric methods are not assumption-free. They relax fixed distributional forms, often gaining flexibility while paying in efficiency, interpretation, or overfitting risk.
Past papers train question-analysis discipline, not fortune-telling. Each problem should return to data type, unknown quantity, statistical tool, calculation path, and contextual conclusion.
Joint PMF problems require listing every cell. Marginalization, conditional probability, and variable transformations are all sums or regroupings of the original cells.
Probability problems are often hard because the viewpoint changes. Define events first, then distinguish conditioning, independence, mutual exclusivity, and Bayes' rule.
A sample is the data, a statistic is a function of the sample, and a sampling distribution is the distribution of that statistic under repeated sampling.
Random variables turn uncertain outcomes into numbers. PMF, PDF, and CDF then let you compute discrete probabilities, continuous interval probabilities, thresholds, and model-score distributions.
A regression table is not a p-value list: coef, SE, t, F, and R-squared answer effect size, uncertainty, single-coefficient tests, overall model signal, and in-sample explanation.
Regularization adds a preference against extreme parameters. Ridge, Lasso, and weight decay trade some training fit for a model that generalizes more reliably.
A reproducible workflow preserves the evidence chain from data to conclusion. Results need data versions, code, seeds, environment, metrics, and raw outputs.
Sampling makes sample statistics fluctuate, and standard error describes that fluctuation. This post separates SD, SE, sampling distributions, and CLT, then connects them to benchmark uncertainty.
A sampling distribution describes how a statistic fluctuates under repeated sampling. Means, proportions, and variances each connect to common distributions used in intervals and tests.
The series does not finish all of statistics. It gives beginners a working map for exams, ML/AI evaluation, causality, Bayesian thinking, time series, and mathematical statistics.
Simple linear regression uses one X to describe the average change in Y. Slope, intercept, residuals, and squared error form the smallest supervised learning model.
Do not start statistics exam prep by memorizing formulas. Start with the sequence of data, probability, sampling, inference, regression, then connect those ideas to model evaluation, A/B testing, and uncertainty in ML/AI.
Time-series data have order. Random splits can leak future information into training and make forecasting or monitoring results look better than they are.
Two-group comparisons start by classifying the outcome and the design: numeric or binary, independent or paired. That choice determines the standard error, test statistic, and conclusion.
GPUtw.ai is a Taiwan-based short-rental GPU cloud. Its main value is not maximum scale, but Taiwan data centers, prepaid credits, Jupyter/ComfyUI/Ollama/vLLM templates, Vault storage, and team billing. Public information is enough for a service introduction, not enough for procurement or production endorsement.
HW0 checks CS181 prerequisites in four problems — y=Xw solvability, optimizing an objective, reasoning about randomness, and OLS in Python. The problem that slows you down most is the gap to patch before HW1.
HW1 uses an 800,000-year ice-core temperature dataset across four problems: kNN and kernel regression, a geometric proof of least squares, basis-function regression, and a probabilistic derivation of ridge and LASSO, ending with a coordinate-descent LASSO implementation.
CS181 2026 is A3 with hw0–6 as the weekly clock (no public recordings); 2025 adds a practical, 2024 has two midterms, 2023 was taught by Weiwei Pan. Start with HW0, then follow hw1→hw6.
Computational biology PhD turned UW-Madison professor Sebastian Raschka launched Ahead of AI on Substack in 2022, publishing monthly deep dives into LLM papers and architectures. Four years later: 200K+ subscribers, zero sponsorships, and a book-newsletter flywheel that proves low frequency and high depth can win in a crowded AI newsletter market.
2021 was the year diffusion models surpassed GANs, self-supervised learning made theoretical breakthroughs, and reinforcement learning confronted weaknesses in its evaluation methodology. NeurIPS received a then-record 9,122 submissions, ICLR’s Score-Based Generative Modeling paper became a theoretical foundation for the diffusion ecosystem, and ICML delivered substantial work on optimization theory and the dynamics of self-supervised learning.
2022 was the year diffusion models took center stage, Chinchilla scaling laws rewrote large-model training, and Chain-of-Thought turned reasoning into an ability that prompts could elicit. NeurIPS passed 10,000 submissions; three of its 13 Outstanding Papers directly concerned diffusion; and Chinchilla and data pruning both challenged the belief that bigger was always better. On the eve of ChatGPT’s release, every required piece fell into place at that year’s conferences.
In 2023, LLMs took over the machine-learning conference agenda. NeurIPS received more than 12,000 submissions; both Outstanding Papers addressed large models, while runner-up DPO became a practical alternative to RLHF within two years. DreamFusion opened the text-to-3D field, ICML spotlighted LLM watermarking and learning-rate adaptation, and the Mamba preprint emerged as the first serious architectural challenger to the Transformer.
ML conference submissions exploded in 2024: NeurIPS received a record 15,671 papers, while ICML and ICLR passed 9,000 and 7,000. Research shifted from training ever-larger models toward spending inference compute more intelligently, making test-time compute scaling the year's defining new direction. VAR beat diffusion with next-scale image prediction, Rectified Flow became the theoretical foundation for Stable Diffusion 3, and ICLR gave its inaugural Test of Time Award to the original VAE paper.
ML conferences broke every submission record in 2025 and pushed peer review to its limit. NeurIPS received 21,575 papers and used more than 20,000 reviewers; ICML passed 12,000 for the first time, and ICLR reached 11,565. Reasoning and agents were the strongest trends. One NeurIPS runner-up, the conference's only perfect-score paper, challenged whether RLVR creates new reasoning ability. Awards for Alibaba Qwen's Gated Attention and a mechanistic theory of neural scaling laws showed a community moving from scaling at all costs toward understanding why scaling works.
ML fundamentals interviews don't test whether you can recite definitions — they test whether you can walk through a structured diagnostic when handed a train/val accuracy gap. Today covers four high-frequency topics: bias-variance decomposition and learning curve interpretation, geometric intuition for L1/L2 regularization and when to pick which, aligning loss functions with business objectives instead of accepting defaults, and why AdamW decouples weight decay from L2 regularization.
Finishing 07-280 means more than reading 24 guides: produce a search engine, supervised-model comparison, CNN/GPT-2 experiments, and a small RL-plus-MCTS system before choosing 07-380, 10-301, or a specialist course.
07-280 is CMU's new Spring 2026 AI+ML core: 24 lectures and 12 main assignments move from heuristic search and CSPs to AlexNet, GPT-2, and AlphaZero. Its public material supports self-study, but complete recordings, Canvas checkpoints, Gradescope, and staff feedback remain unavailable.
Lecture 1 uses an alien autoencoder, the scope of AI and ML, and AI history to establish the course's coordinate system: an intelligent system turns inputs into representations and decisions under uncertainty.
Lecture 5 formulates machine learning through `X → Y`, loss, risk, and empirical risk minimization: a training set only gives average observed loss, while the real objective remains generalization over an unknown distribution.
Lecture 6 recursively grows a tree from decision stumps, measures label uncertainty with entropy, and selects splits by `I(Y;W)=H(Y)-H(Y|W)`; this is computationally practical greedy ERM, not a global optimal-tree guarantee.
Lecture 7 applies ERM to linear functions and squared loss, moves from a one-dimensional slope to `argmin ||y-Xθ||²`, and derives the normal equation when `XᵀX` is invertible.
Lecture 8 moves from a one-dimensional parabola to vector gradients and compares batch GD, SGD, and mini-batches; the learning rate determines whether updates converge, oscillate, or diverge.
Lecture 9 models P(y=1|x) with a sigmoid instead of directly predicting 0 or 1, learns parameters with cross-entropy and convex optimization, and extends naturally to softmax regression.
Lecture 10 uses φ(x) to let linear models express nonlinear functions, then controls the resulting overfitting with train/validation/test separation, L1/L2 regularization, and model selection.
Lecture 11 expands a logistic unit into a multilayer network: linear layers produce z, activations produce a, and multiple neurons jointly learn a feature transform trained through a final loss.
Lecture 12 treats a network as a computation graph: the forward pass stores intermediates, the backward pass propagates upstream gradients, and local linear, activation, and softmax rules compute every parameter gradient efficiently.
Lecture 16 starts from likelihood p(D|θ), uses i.i.d. to factor the joint probability and logs to turn products into sums; Bernoulli MLE yields sample proportions, conditional Bernoulli yields logistic cross-entropy, and Gaussian noise yields squared error.
Lectures 1–12 form one decision pipeline: define states, moves, and objectives, then use heuristics, losses, regularization, and backpropagation to control an otherwise intractable search space.
DSPy replaces handwritten prompt strings with task Signatures, execution Modules, and Optimizers that compile better instructions and examples against a dataset and metric.
Hugging Face Hub is a collaboration layer for versioned models, datasets, and applications. Datasets handles data, Spaces runs demos, while Inference Providers and Endpoints provide managed inference.
Linear regression is more than a best-fit line: Chapter 1 connects squared loss to gradient descent, normal equations, maximum likelihood, and locally weighted regression.
Chapter 2 derives logistic loss from a sigmoid probability model, then contrasts it with the perceptron and extends it through softmax and Newton's method.
Chapter 3 uses exponential families, natural parameters, and link functions to place least squares and logistic regression inside one modeling template.
Chapter 5 replaces high-dimensional feature inner products with kernels, letting inner-product-based linear algorithms learn nonlinear functions without constructing the features.
Chapter 7 decomposes neural networks into composable modules and uses backpropagation and vectorization to explain how deep models can be trained efficiently.
Chapter 8 decomposes test MSE into irreducible noise, squared bias, and variance, then uses uniform convergence and VC dimension to explain when training performance transfers to new data. Double descent shows why parameter count is not a universal measure of complexity.
Chapter 9 presents three controls on generalization: explicit complexity penalties, optimizer-induced implicit regularization, and model selection on data excluded from training. MAP estimation then connects a Gaussian prior to an L2 penalty.
Chapter 10 introduces unsupervised learning through k-means: alternating updates make distortion non-increasing and numerically convergent, but do not guarantee a global optimum.
Chapter 11 starts from soft assignments in Gaussian mixtures, uses Jensen's inequality to construct the ELBO, interprets EM as alternating maximization over a variational distribution and model parameters, and extends the idea to VAEs through approximate posteriors and reparameterization.
Chapter 12 formulates PCA as geometric optimization: maximize projected variance along a unit direction to obtain the leading eigenvector of the covariance matrix. The top k eigenvectors give both maximum retained variance and minimum linear reconstruction error.
Chapter 13 models ICA as x=As: observations are unknown linear mixtures, and the goal is to estimate W=A^{-1} to recover independent, non-Gaussian sources. A Jacobian determinant enters the transformed density and leads to the Bell–Sejnowski likelihood update.
Chapter 14 starts with a fixed Gaussian noising Markov chain and learns to reverse each transition. The ELBO turns reverse-kernel matching into weighted noise prediction, while the continuous-time view explains reverse drift through the score ∇log p_t.
Lectures 19–25 connect rational decisions and VPI to machine learning, while Project 5 uses PyTorch for regression, classification, CNNs, attention, and an optional character-GPT.
CS189 exists online in several terms. From post 2 onward this series uses Spring 2026 (eecs189.org/sp26, A3) as its base, walking through each lecture block and assignment; Spring 2025 (Shewchuk) is a different classic route with 25 lectures of notes, HW1–7 and past exams, but its official recordings sit behind a bCourses login; Fall 2026 is still in progress and serves only as a reference. This post is the series entry point and full table of contents.
07-380 Fall 2026 is the first offering of CMU's new AI II: 26 lectures from logic, planning and optimization to probabilistic graphs and generative systems. As of the 2026-09-29 course site, the Lec1–9 slides, PR1–6 notes, Rec1–5 (with solutions) and HW1–3 are public, and this site now has 14 lecture-by-lecture guides for them. Slides after Lec10, HW4–7 and the final project are not out yet, so the course as a whole is still A2.
HW6 combines generalization, MLE/MAP, probabilistic learning, fairness metrics, and social impact in one written assignment about assumptions and tradeoffs.
The final written assignment combines ensembles, clustering, representation, and recommendation to test whether you can choose a learning paradigm from problem structure.
Spring 2026 publishes material for 27 lectures and nine homework bundles; outsiders can do the core work but cannot access Panopto, Piazza, Gradescope, or official homework solutions.
In 2026, CMU recombined its separate general-AI and SCS machine-learning introductions into the 07-280 → 07-380 sequence. This is a redistribution of content and prerequisites, not a pair of simple course renames.
CS50 AI is Harvard's most complete public entry point, but the Summer 2026 course still uses 2020 recordings and assignment assets while the rolling OCW projects have moved to other editions. CS181 Spring 2026 exposes current homework and notes without current recordings; CS182 Fall 2026 has not yet completed an offering.
Logistic regression turns a linear score into a Bernoulli probability with sigmoid; the gradient xⱼ(y-ŷ) follows directly from the log-likelihood chain rule.
Lambda Cloud provides on-demand GPU VMs and 1-Click Clusters; it offers direct AI compute environments rather than automatically solving training, serving, and MLOps.
Replicate abstracts GPUs behind versioned models, predictions, Cog, and deployments; integrators still own version pinning, async workflows, webhook verification, data persistence, and spending limits.
RunPod Pods fit interactive and persistent GPU work, while Serverless fits queued or load-balanced inference; choosing incorrectly mixes persistence, cold starts, and retry semantics.
The three things you need to self-study CS229 run on three different clocks. The lecture notes are 278 pages and were recompiled in August 2026. The newest problem sets you can download are from summer 2020. The self-assessment Stanford Online tells you to attempt before enrolling is a PDF created in 2008. Seventeen lectures from spring 2026 are public, and the last three are mislabeled.
Berkeley has no standalone undergraduate AI degree. A workable path builds on the CS BA or EECS BS foundation, enters through either CS188's broad AI curriculum or CS189's mathematical machine learning curriculum, then branches into deep learning, NLP, vision, or reinforcement learning. Many 2025–2026 courses are A3, but the newest class, the newest stable URL, and the best self-study edition are not always the same.
CMU's current BSAI now runs through 07-280 and 07-380 before branching into an NLP/vision core and four AI clusters. 07-380 debuted in Fall 2026, and the same semester launched a graduate-level 11-768 AI Agents course. The residual Spring 2026 materials for 07-280 and the complete 10-301/601 site already support self-study; retired 15-281 remains a useful legacy route.
MIT has offered Course 6-4, a formal BS in Artificial Intelligence and Decision Making, since 2022. For an outside learner, however, the current degree requirements, the 2025–2026 course sites, and the best OCW editions rarely line up. A workable route follows 6-4's programming, algorithms, linear algebra, and probability foundation, then selects among 6.S191, 6.3900, 6.4110, 6.7960, vision, and robotics according to what is actually public.
Every lecture in CS109's Summer 2026 offering ships with an official LLM Learning Guide — six concepts, a Learn prompt and a Test me prompt for each, written week by week across the quarter for a total of 23 PDFs. The same course's honor code Rule 4 forbids asking an LLM to solve your homework, and 65% of the grade sits in proctored exam rooms. Those two facts are halves of one design.
The finale recaps the CS161 toolbox and points toward LP duality, Reed–Solomon coding, and ML-assisted algorithms. Officially, this lecture has slides but no notes.
ML fundamentals interviews don't test formula memorization — they test whether you can explain concepts intuitively and hold up under follow-up questions. High-frequency topics: the practical meaning of bias-variance tradeoff, the selection logic for L1/L2 regularization, why cross-entropy beats MSE for classification, SGD vs. Adam tradeoffs, and how precision/recall priorities differ by scenario.
AI Engineer interviews go beyond ML — big tech emphasizes system design and coding, startups look for end-to-end delivery, and AI-native companies test LLM engineering depth. Strategy: identify your target company types first, then allocate prep time across six dimensions (ML fundamentals, system design, LLM applications, coding, paper reading, and behavioral).
Google's Professional ML Engineer exam guide was rewritten in 2026: Vertex AI is renamed Gemini Enterprise Agent Platform throughout, so older study material no longer matches the product names in the questions. This guide uses the official six-section weighting as its skeleton, listing what each section tests, which official materials cover it, and what to build — plus a study schedule whose reasoning is spelled out. Official specs: $200, two hours, 50–60 multiple-choice and multiple-select questions, two-year validity, 3+ years of industry experience recommended including 1+ year on Google Cloud.
Andrew Ng demonstrates error analysis on a deep researcher: columns are the pipeline stages, rows are 10 to 100 queries, you only look at the ones that went badly, and you mark each cell where something broke. The percentages don't have to sum to 100%. He says it takes three or four hours and saves weeks of going the wrong direction — and the fraction of people who actually do it is far below 100%.
Andrew Ng walks a face-recognition door system through the entire project lifecycle, and the whole lecture has one thesis: speed. He gives teams a two-day deadline, on the reasoning that 'time spent preparing data should be commensurate with the time it takes to train the model once.' It closes on a line: my job is to build something that actually works, and that is not the same as building something that works on the test set.
ArduPilot's ModeReason enum is the exhaustive list of on-board autonomous decisions: of 56 values, 43 are the aircraft deciding for itself, 21 of those because it detected something dangerous, and exactly one — SOARING_THERMAL_DETECTED — because it found an opportunity. As for end-to-end, PX4 mainline's mc_nn_control has a 10 KB tensor arena and a Kconfig default of n; ArduPilot mainline has none.