Skip to content

CMU 10-423 L15–L16: Scaling Laws and Mixture of Experts — How Big Should the Model Be, and How Do You Compute Only Part of It?

Sep 30, 20261 min
TL;DRThe first two lectures of the Scaling Up unit in CMU 10-423 Spring 2026 answer two questions. The second half of L15 covers scaling laws: Kaplan 2020 says 8x more parameters needs only about 5x more data, Chinchilla says scale both equally, and the Phi models and data-filtering scaling laws add data quality as a third axis. L16 covers MoE: feed-forward layers hold most of GPT-3's parameters, so split them into experts and send each token through only the top k. Memory follows total parameters, compute follows active parameters, and the price is load balancing and training stability. No programming homework covers this half of the course; quizzes, practice exam question 13, and the final project do.

🌏 中文版

This guide is based on the Spring 2026 edition of CMU 10-423/623/723 Generative AI. It is part 16 of Reading CMU 10-423 and opens the fifth unit, "Scaling Up". The previous post, HW4, was the last programming assignment. From here on the course checks your learning through quizzes, HW623 (10-623/723 only), and the final project.

Official materials used: the Lecture 15 slides (37 pages; the Querying Transformer half is covered in part 14, so this post reads only the Scaling Laws section), the Lecture 16 slides and their inked version, the course schedule, and the practice exam on the Coursework page. I downloaded and checked all of them on 2026-09-30. The schedule lists no readings for these two lectures, so every paper cited here is one the slides cite.

Version note: Both PDFs linked from the March 16 Lecture 16 row have a cover reading "Matt Gormley & Pat Virtue, Mar. 17, 2025", a reminder slide with Spring 2025 HW4 dates, and a 2025 creation date. Spring 2026 reused last year's MoE deck. The Scaling section of L15 is marked "Scaling slides credit: Pat Virtue". Recordings are on CMU's internal Panopto and not visible from outside, so this post relies entirely on the slides.

Why this comes after multimodal models

The schedule groups L15–L18 as "Scaling Up". The first 14 lectures asked what models look like and how to train them. This unit turns to engineering: with limited money and GPUs, how big should the model be, how much data should it see, and what do you do when it doesn't fit?

The four lectures split the work cleanly. L15–L16 answer "how big" and "can a big model compute less", and L17–L18 in the next post answer "how do you actually run it".

L15, second half: scaling laws

It starts with a question

The Scaling section opens with two timelines, one for language models and one for image generation, and a question: Transformers appeared in 2017 and took over NLP immediately, so why did Vision Transformers take until 2021? The slides don't answer directly. They move to a table of LLM sizes from GPT-2 to LLaMA-3, with the section's guiding question beside it: how did Meta choose this combination of training tokens and model parameters?

Some rows from the table (as on the slide):

ModelYearTraining tokensParameters
GPT-22019~10 billion1.5 billion
GPT-32020300 billion175 billion
Chinchilla20221.4 trillion70 billion
LLaMA-220232 trillion70 billion
LLaMA-3202415 trillion405 billion

Next comes a "how much did it cost to train Llama?" exercise. It gives a GPU price (around $15k), cloud GPUs at $1–4 per hour, and the electricity cost of a 700W card, and asks you to estimate the cost of Llama-3 70B. The answer box is empty on the slide, and the schedule has no inked version of L15.

Power laws and Kaplan 2020

Most LLM scaling laws assume a power law between loss and some quantity, of the form $f(x) = c,x^{-k}$. The slide plots the same curve on linear and log-log axes: on log-log axes it's a straight line.

For Kaplan et al. 2020, the slides list what the experiments varied: parameters N (768 to 1.5B), data D (22M to 23B tokens), compute C, model shape (depth, width, heads), context length (up to 1024), and batch size. They then measured each model's test loss.

The seven takeaways the slides step through:

  1. Three quantities dominate: parameters N, tokens D, and FLOPs C.
  2. Model shape doesn't matter much.
  3. Performance improves as long as N and D grow together.
  4. Training and test loss curves follow predictable power laws.
  5. Larger models are more sample-efficient.
  6. You don't need to train to convergence to get good performance.
  7. The best batch size also follows a power law, and it's huge (1–2M tokens).

The last of these slides quotes Kaplan: every time model size grows 8x, data needs to grow only about 5x to avoid a penalty. The next line reads: "But Hoffman et al. (2022) tell a very different story!"

Chinchilla: everyone was using too little data

Hoffmann et al. 2022 (Chinchilla) fixed the compute budget C, varied both D and N, measured L(N, D), fit a model that predicts loss for any N and D, and used it to find the optimal model size.

The slides boil it down to one line: everyone had been using far too little data. Kaplan paired 8x parameters with 5x tokens. Chinchilla says scale them proportionally (2x parameters, 2x tokens). Under the same compute budget, a smaller model trained on more data does much better. The slides don't write out Meta's answer, but the guiding question is meant to be read through this result. In the same table, Chinchilla pairs 70B parameters with 1.4T tokens, far smaller and far more data than GPT-3's 175B parameters on 300B tokens.

A third axis: data quality

The last part is titled "Adjusting quality of data":

  • The Phi models: instead of growing the model or the data, improve the data. Phi-1 ("Textbooks Are All You Need") matched much larger models on coding, and Phi-1.5 extended the result to general LLMs.
  • Scaling laws for data filtering: Goyal et al. 2024 deal with trading off quantity, quality, and compute. The slides' summary: the more compute you have, the less you need to filter.
How the practice exam tests this section

Question 13, "Scaling Laws" (4 points), on the Spring 2026 practice exam asks about the link between compute budget and data filtering, what the Hoffmann study found about model size versus data, how test loss changes with parameter count, why simply growing the model gives diminishing returns, and the key insight behind the Phi models. It also has two MoE short answers: why MoE is more efficient and why it is hard to train. Solutions are in a separate PDF. Write your own answers first.

The L18 reminder slide says the March 30 evening exam covers Lectures 1–15, so scaling laws are on the exam and MoE is not. The schedule puts Quiz 5 over L16–L20.

L16: Mixture of Experts

The parameters live in the feed-forward layers

L16 opens with a GPT-3 parameter breakdown, taken from the OLMoE paper. Of 174.57B parameters, feed-forward layers hold 115.97B, attention 57.99B, and embeddings only 0.62B. The slides ask you to work out how each layer type's count comes about.

The point: more than half of a Transformer LLM's parameters sit in the feed-forward layers, so that's where to cut.

Splitting a layer into experts

The slides explain "experts" in two steps:

  1. A linear layer $z = Wx + b$ can be split by rows into three pieces $z_i = W_i x + b_i$ with the same total parameters.
  2. A feed-forward layer $y = U\sigma(Wx + b) + c$ can likewise be split into three small feed-forward networks. The two are equivalent when $W$, $b$, and $U^T$ are the three pieces stacked. The slides leave what $c$ should be as a fill-in-the-blank.

Note the slide's caveat: in MoE, each expert is not a linear layer but a feed-forward network with one hidden layer.

Dense and sparse gating

The MoE output is a weighted sum $y = \sum_{i=1}^{N_e} G(x)_i E_i(x)$. What changes is the gate $G$:

  • Dense MoE: $G(x) = \text{softmax}(x \cdot W_g)$, so every expert gets a nonzero weight.
  • Sparse MoE: $G(x) = \text{softmax}(\text{topk}(x \cdot W_g + b_g, k))$ keeps only the k highest scores. The slides note that sparsely-gated MoE was first proposed for RNNs but applies generally and is now popular in Transformers.
  • Noisy top-k: add Gaussian noise to the gate scores, scaled by an input-dependent term.
  • Mixtral: the same top-k gating, with each expert replaced by a SwiGLU feed-forward network (Mixtral paper).

One initialization detail: $W_g$ and $W_{noise}$ start at all zeros, which gives no signal at first and only a little noise.

Put experts on different GPUs and traffic jams follow

Expert parallelism puts each expert on a different GPU and routes each token to k experts. The slides use two small examples. With 3 devices, each able to hold 3 tokens, 6 tokens, and k = 1, some devices fill up while others sit idle. With 4 tokens and k = 2, token 3 gets routed to a device that is already full and can't fit.

Left alone, the gate concentrates on a few experts that happened to be popular early in training. The slides give OLMoE's fix, noting that many variants exist: add two regularizers to the loss.

$$L = L_{CE} + \alpha L_{LB} + \beta L_{RZ}$$

  • Load balance term $L_{LB} = N_e \sum_i f_i P_i$, where $f_i$ is the fraction of tokens in the batch routed to expert i and $P_i$ is the probability assigned to expert i. It pushes load to spread out.
  • Router z-loss $L_{RZ}$ penalizes large router logits to stabilize training.

Active parameters: memory and compute are counted separately

The k in top-k is usually small. The slides' two examples: Mixtral uses k = 2 with $N_e$ = 8, and OLMoE uses k = 8 with $N_e$ = 64.

Active parameters are the ones the router selects for computation. Roughly:

  • GPU memory ∝ total parameters
  • FLOPs ∝ active parameters

That's the MoE tradeoff: spend more memory to compute less per token. The slides follow with a Mixtral vs. Llama-2 comparison, OLMoE's hyperparameter table, and the "performance vs. cost" plots from the OLMoE paper, concluding that MoE offers a good tradeoff between performance and FLOPs.

How many experts?

The last two slides contrast two eras. Early MoE work on LSTM language models favored a very large number of experts, while recent Transformer LMs favor comparatively few. The closing slide shows the routed-LM scaling laws of Clark et al. 2022, tying the lecture back to L15.

How to study these two lectures

  1. In L15, read only the Scaling section (from slide number 13 on). Put Kaplan's seven takeaways next to Chinchilla's one line and write down their different answers to "how much more data for 8x the parameters?"
  2. Work out GPT-3's per-layer parameter counts yourself (the exercise on L16 slide number 8) and see why feed-forward layers dominate.
  3. Use the two expert-parallelism examples on L16 slide numbers 21–22 and work out by hand which tokens get dropped.
  4. Finish with practice exam question 13, then check the solutions.

What to do: tonight, open L16 slide number 12 and fill in what $c$ should be when a feed-forward layer is split into three experts. If you can, you've understood that experts are just a big feed-forward layer regrouped.

What this post can and can't confirm

Confirmed: the text, formulas, and citations on both slide decks; schedule dates and quiz coverage; the exam coverage on the L18 reminder slide; and the question stems in practice exam question 13. Not confirmed: what was said or drawn in class (recordings are on Panopto, and L15 has no inked version), the official answer to the Llama cost exercise, and numbers that appear only in figures the text layer doesn't capture (such as the Mixtral vs. Llama-2 comparison).

Further reading

Series: Previous: HW4: Text-to-Image with a Q-Former | Next: L17–L18: Distributed Training, FlashAttention, and Efficient Decoding | Series overview

References