Skip to content

AI Engineer Interview Daily — 2026-10-05: ML Fundamentals

Oct 5, 20261 min
TL;DRToday's ML Fundamentals rotation covers four core concepts: the efficiency gap between Grid Search, Random Search, and Bayesian Optimization (Bayesian uses a surrogate model that remembers past trials, roughly 7x fewer iterations and 5x faster), how Huber Loss uses a single threshold delta to stitch MSE's smooth gradient together with MAE's outlier robustness, how the curse of dimensionality breaks distance metrics in high-dimensional space, and what PCA's role looks like one layer beneath 'dimensionality reduction tool.' The practice question comes from a real Pinterest Machine Learning Engineer interview (PracHub's 2026 question bank): explain what causes vanishing gradients in deep neural networks, then discuss the specific mathematical trade-off each of ReLU, Leaky ReLU, and ELU makes to mitigate it — the breakdown threads 'why the chain-rule product collapses toward zero' through to 'which new problem you trade it for' into one complete reasoning chain.

🌏 中文版

Today's Topic

Monday's rotation is ML Fundamentals, and today's four concepts plus the practice question all sit on the same thread: where training can go wrong, and what mathematical tool fixes it. Hyperparameter search determines whether you find a good configuration efficiently at all; Huber Loss determines whether outliers can drag a regression model off course; the curse of dimensionality determines whether your distance metric still means anything in high-dimensional space; vanishing gradients determine whether deep network training works at all. These are exactly the kind of "why" follow-ups that show up from phone screens through onsite rounds — reciting definitions doesn't get you through; interviewers want the mechanism and the trade-off.

Core Concepts Cheat Sheet

Grid Search, Random Search, and Bayesian Optimization split on whether they remember past trials

Grid Search exhaustively tries every combination on a fixed grid — continuous hyperparameters get discretized into fixed steps, and it has zero memory of previous trials. It's the slowest but most exhaustive baseline. Random Search samples combinations at random instead, dropping the grid's discretization limit and running faster, but it's still fundamentally "blind guessing" — it doesn't adjust where to look next based on what it's already tried. Bayesian Optimization's key difference is that it fits a surrogate model (typically a Gaussian Process) to predict how good each configuration is likely to be, then uses an acquisition function to trade off between exploiting known-good regions and exploring unknown ones — and the surrogate's predictions sharpen with every round. The efficiency gap in practice is concrete: a commonly cited figure is roughly 7x fewer iterations and 5x faster convergence for Bayesian Optimization versus Grid Search. Being able to name the "surrogate model plus acquisition function" mechanism shows a lot more depth than reciting three names.

Huber Loss stitches MSE's smoothness to MAE's robustness with a single threshold

MSE (mean squared error) squares the error, which amplifies outliers — the gradient becomes very sensitive to large errors, and training can get pulled off course by a handful of them. MAE (mean absolute error) treats every error equally, making it robust to outliers, but its gradient is discontinuous near zero, which hurts convergence to a precise solution. Huber Loss sets a threshold delta: below it, errors are squared (keeping MSE's smooth convergence on small errors); above it, errors switch to linear (keeping MAE's resistance to outlier amplification). A common follow-up is how to choose delta: a smaller delta pushes the loss closer to MAE — more robust to outliers but slower to converge; a larger delta pushes it closer to MSE — faster convergence but outliers carry more weight. In practice, people often set a starting point from some quantile of the residual distribution (around the median-scale region) and tune it against a validation set.

The curse of dimensionality isn't "more information" — it's distance metrics breaking down first

As feature dimensionality grows, data points become increasingly sparse — maintaining the same data density requires a sample size that grows exponentially with dimension. That means any method relying on "close enough neighbors" to mean something — KNN, density-based clustering — breaks down first in high dimensions: the relative gap between the nearest and farthest neighbor shrinks, and distance itself loses discriminative power. That's why the curse of dimensionality isn't just a "too many features, runs slower" performance problem — it's the model's underlying assumption failing. PCA is often framed as a dimensionality-reduction tool that mitigates this, but the deeper answer for an interview is that PCA isn't simply discarding features — it finds the orthogonal directions of maximum variance in the data, preserving the most information in fewer dimensions, which incidentally restores distance's discriminative power in the lower-dimensional projection.

Vanishing gradients: the chain-rule product in backpropagation shrinks exponentially in deep networks

When backpropagation computes the gradient for a given layer's weights, it multiplies that layer's gradient by the derivative of every activation function in front of it. Saturating activations like sigmoid and tanh have derivatives that approach zero when the input is very large or very small, and the deeper the network, the more of these sub-1 derivatives get multiplied together — so the gradient reaching shallow layers shrinks exponentially toward zero. That means the shallow layers barely learn anything, and training stalls. This exact mechanism is the core of today's practice question, since the choice of activation function directly determines whether this chain-rule product collapses.

Today's Practice Question

The Question

Vanishing Gradient Problem: What causes gradients to vanish in deep neural networks? Discuss how different activation functions (such as ReLU, Leaky ReLU, and ELU) mitigate this issue, and explain their mathematical trade-offs.

Source: Pinterest Machine Learning Engineer interview question bank (PracHub 2026) Difficulty: Intermediate Round: Technical Phone Screen / Virtual Onsite ML theory round

How to Break It Down

  1. Clarify first: Confirm the network depth and the activation function originally in play (the default assumption is usually a deep sigmoid/tanh network) — that determines whether you volunteer the "what new problem shows up after switching activations" extension.
  2. Build a framework: Start with the cause (chain-rule product plus saturating activation derivatives approaching zero), then cover how different activation functions mitigate it, and only then move to the auxiliary techniques beyond activation functions.
  3. Go deep on the core: Sigmoid and tanh derivatives approach zero in their saturation regions, and in a deep network these sub-1 derivatives multiply together, shrinking the gradient exponentially. ReLU's derivative is a constant 1 in the positive region, which solves the saturation problem — but trades it for a new risk, Dying ReLU: once a neuron's weighted input stays negative, its output and gradient are permanently zero, and the neuron effectively dies and never updates again. Leaky ReLU gives the negative region a small fixed slope (say, 0.01), keeping the gradient from hitting exactly zero, at almost no extra computational cost. ELU uses an exponential function in the negative region so the output smoothly asymptotes toward -α; beyond avoiding a zero gradient, it also pushes the output mean closer to 0 (which helps convergence), but the exponential computation costs more than Leaky ReLU's.
  4. Close strong: Summarize it as one line — "ReLU solves saturation but introduces dying neurons; Leaky ReLU and ELU both redesign the negative region for a different trade-off: Leaky ReLU trades for the cheapest possible gradient floor, ELU trades for a smoother output distribution at a higher computational cost" — then add that modern deep networks rarely rely on the activation function alone, pairing it with residual connections (giving gradients a shortcut past saturating layers) and an initialization scheme like He initialization, designed specifically to match ReLU-family activations.

Sample Answer (What You'd Actually Say)

Start with the cause: Vanishing gradients come from the chain rule in backpropagation — computing a shallow layer's gradient means multiplying the current layer's gradient by the derivative of every activation function after it, all the way back. Sigmoid and tanh saturate when their input is large or small in magnitude, and their derivative approaches zero there. The deeper the network, the more of these sub-1 numbers get multiplied together, so the gradient reaching shallow layers shrinks exponentially toward nothing — which is exactly why early deep networks struggled to train past a handful of layers.

What you trade when you switch activations: ReLU has a constant derivative of 1 in the positive region, which directly fixes the saturation-driven gradient collapse — one of the key reasons deep networks could later stack dozens of layers. But ReLU isn't free: if a neuron's weighted input stays negative, its output and gradient both get stuck at zero permanently — that's Dying ReLU. Leaky ReLU's fix is a small fixed slope in the negative region, guaranteeing a nonzero gradient at almost no extra cost. ELU instead uses a smooth exponential transition in the negative region — its advantage is an output mean closer to 0, which helps convergence, but the exponential computation costs more than Leaky ReLU's. It's a straight trade of compute for smoothness and faster convergence.

A more complete answer than just the activation function: I'd add that modern deep networks don't rely on swapping activation functions alone — they pair it with residual connections, so gradients can flow back to shallow layers through a shortcut instead of depending entirely on the multiplicative chain, plus weight initialization matched to the activation choice — He initialization is specifically designed for the ReLU family, and assumes different activation properties than Xavier initialization does. Swapping the activation function alone treats the symptom; all three together are what makes deep network training stable today.

Self-Check List

Use this table to check whether your answer covered the key points:

Check itemCovered?
Clearly explained the cause: chain-rule product plus saturating activations with derivatives near zero
Pointed out that ReLU solves saturation but introduces Dying ReLU as a new problem
Explained Leaky ReLU's and ELU's negative-region designs and their respective mathematical trade-offs
Mentioned the computational-cost dimension of the trade-off (Leaky ReLU low, ELU high)
Extended beyond the activation function to auxiliary techniques (residual connections, matching weight initialization)
Bonus: noted that He initialization and Xavier initialization assume different activation properties

Further Reading

References