🌏 中文版
This guide follows the Spring 2025 offering (NCCU term 1132) of Yen-Lung Tsai's "Generative AI: Text and Image Synthesis Principles and Practice" at National Chengchi University. It is part 5 of the Reading NCCU Yen-Lung Tsai Generative AI series and follows L04 LLMs Are Simpler Than You Think. The course is taught in Mandarin.
Two official sources back this post: the Lecture 5 recording (2025-03-18, 3 h 3 min) and the slide deck GenAI05 The Mathematics of Transformers (67 pages; the cover reads "The Mathematics of RNNs and Transformers"). On the Chang Gung satellite class page this week is titled "Transformers 全攻略" ("The Complete Guide to Transformers") and the homework column says "no homework." Access level: A3.
If you'd rather skip the math: this is the steepest stretch of the series. Read just the first paragraph of each section and skip the collapsed blocks, or jump straight to L06 LLM Applications and Ethics. The application lectures that follow won't leave you stuck.
Where this week fits
L04 framed an LLM as a machine that continues text by guessing the next word, and gave Q/K/V a brief cameo. L05 goes back and finishes the job: every component of a Transformer, taken apart, is matrix multiplication.
The recording's structure: the first hour reviews RNNs, fills in linear algebra, and derives attention and √d_k. The second covers multi-head attention, encoder/decoder, masking, positional encoding, residuals, and normalization. The third has two student lightning talks (crop price forecasting and a tarot-reading app), and the TA session is devoted to "the mysterious √d_k inside softmax."
Writing an RNN as matrices
Slides 3–8 use a recurrent layer with two inputs and two RNN neurons. An RNN neuron looks like a fully connected layer; the only difference is that its previous output h_{t−1} comes back in and joins the current input in deciding the current output.
Look closely and an RNN neuron is an ordinary neuron with two kinds of input: the "normal" input x_t and the previous hidden states h_{t−1}. Multiply each by a weight matrix and every neuron is computed at once. This "compute a whole batch at once" notation is the key to reading the Transformer.
RNN neurons in matrix form (slides 5–8)
h_t = σ( W_Xᵀ x_t + W_Hᵀ h_{t−1} + b )
Neuron 1 expanded:
h_t¹ = σ( w₁₁ˣ x_t¹ + w₂₁ˣ x_t² + w₁₁ʰ h_{t−1}¹ + w₂₁ʰ h_{t−1}² + b₁ )
Linear algebra 101: two things to remember
Slides 10–14 teach exactly two points, and around minute 24 the recording labels them "the two quick linear-algebra essentials":
- Matrix multiplication goes row-then-column. For C = AB, c_ij is the dot product of row i of A and column j of B. The dot product is the core operation.
- A row vector times a matrix is a linear combination of the matrix's rows. Textbooks usually write Ax with column vectors, but Google loves row vectors, so the Transformer paper writes xA: each component of x is the weight on one row of A.
The second point is the most important technique in the lecture. Attention's "weighted average" later on is exactly this linear combination.
The two rules as formulas
Dot product: u = [a₁, …, a_p], v = [b₁, …, b_p]
⟨u, v⟩ = u vᵀ = a₁b₁ + a₂b₂ + … + a_p b_p
Row combination: x = [x₁, …, x_m], rows of A are a₁, …, a_m
x A = x₁ a₁ + x₂ a₂ + … + x_m a_m
Q/K/V: use a question to find the best-fitting answer
In 2017 Google published Attention Is All You Need, introducing the Transformer, originally meant to replace RNNs. The slides remark that the title shows Google was aiming even higher.
The intuition on slides 17–20: you have a sequence of information x₁, …, x_T (say, a passage of text), and each x_t has two representation vectors, k_t (key) and v_t (value). A query q (a "question") arrives, and you want the representation h that best fits q given everything before it.
The recipe:
- Compute the relevance between q and each x_t. Google chose the simplest option, the dot product q·k_t. The slides stress that relevance can "basically be computed any reasonable way."
- Softmax all the relevance scores into weights α₁, …, α_T.
- h = α₁v₁ + α₂v₂ + … + α_T v_T. That's the row-vector linear combination from the previous section.
Self-attention means each x_t also produces its own q_t, and every position takes a turn as the query. Stack all the q's into a matrix Q and you get the formula Google is so proud of.
Why divide by √d_k
Slides 31–33 put it bluntly: using dot products for attention strength actually worked worse than alternatives (such as training a neuron to compute it), mainly because softmax is winner-take-all and inflates the top weights far beyond reasonable values.
The example: five scores 3.9, 3.2, 1, 0.3, 1.1 give 61%, 30%, … under plain softmax. Two scores that were close end up a factor of two apart. Divide everything by τ = √5 and they become 1.74, 1.43, 0.45, 0.13, 0.49, which softmax turns into 40%, 29%, 11%, 8%, 11%. Pull the numbers back toward 0 and the problem is solved.
The slides' view: any suitably large number works here, say √9487, and writing √d_k "is basically just to make it look sophisticated." τ can also be treated as a tunable hyperparameter. This week's TA session (recording from 2:24:15) covers the same square root.
Deriving attention (slides 21–34)
Single query:
e_t = ⟨q, k_t⟩ = q k_tᵀ
stack k_t and v_t into matrices K and V (one vector per row)
[e₁ … e_T] = q Kᵀ
h = softmax(q Kᵀ) V ← the weights linearly combine the rows of V
All queries at once:
Attention(Q, K, V) = softmax(Q Kᵀ / √d_k) V, d_k = dimension of the key vectors
Where the vectors come from (all W are learned; x_t is a word embedding):
q_t = x_t W_Q, k_t = x_t W_K, v_t = x_t W_V
Multi-head, encoder, decoder, and masking
Multi-head attention (slide 35): there's no reason attention should come in only one flavor, so define the n-th attention with its own learned matrices, Attention(Q W_n, K W_n, V W_n) (the slides' shorthand; in practice Q, K, and V each get their own matrix), and run several in parallel.
The encoder and decoder differ (slides 36–39):
- In the encoder, q, k, and v all come from the input vectors themselves: self-attention.
- The decoder has two attention layers. The lower one is masked multi-head self-attention. The middle one is the only layer that is not self-attention: k and v come from the encoder, q comes from the decoder.
- A Transformer outputs as many vectors as it receives words (word vectors). The decoder works the same way, except positions not yet generated are masked, so it can't peek at later words.
The second hour of the recording also has a segment on "why memory limits the number of words" (from 1:10:22) with no matching slide; this post doesn't expand on it.
Positional encoding: marking order with clocks running at different speeds
An RNN genuinely reads one word at a time. A Transformer, to parallelize, runs self-attention in one shot, so word order isn't actually taken into account. Position information has to be added to the original word embeddings; that's position encoding (slide 41).
The intuition on slides 42–43 is elegant. Look at any positional number system, such as the decimal number 9487. The ones digit changes every step, the tens digit every 10, the hundreds every 100: the higher the digit, the longer its period and the lower its frequency.
We want a "fantasy number system" that is continuous, periodic, and small in magnitude (so it doesn't drown out the embedding). Sine and cosine fit. So the embedding's dimensions are grouped in pairs that share a frequency, with low-order dimensions at high frequency and high-order dimensions at low frequency.
Why use sin and cos together? Slide 48's answer: one number per digit isn't enough, and (sin, cos) is a point on the unit circle, like the hand of a clock. Low-order clocks spin fast; high-order clocks spin slowly.
The recording also stresses (at 1:28:16) that positional embeddings aren't limited to this one method; sin/cos is just the original paper's choice.
The original paper's position encoding (slides 44–49)
The position vector for word t, p_t = [p₀, p₁, …, p_{d−1}], is added to embedding x_t
ω_k = 1 / 10000^{2k/d}
p_{2k} = sin(ω_k · t)
p_{2k+1} = cos(ω_k · t)
The slides flag two details. The original paper usually indexes from 1, but here indexing must start at 0 so that frequencies keep decreasing, ω₀ > ω₁ > …. And polar coordinates put cos first, yet (sin θ, cos θ) also traces the unit circle, with θ = 0 pointing to 12 o'clock and moving clockwise. The slides add, "though we don't know what Google had in mind."
Keeping Transformers stable: residual connections and normalization
Back at the original architecture diagram, the piece not yet explained is Add & Norm: residual connections and layer normalization. The slides call these the best techniques of their day for making networks deeper and more stable, and say that's still roughly true.
ResNet-style residuals: learn only what's missing
ResNet changes a layer's output from ℓ(z) to z + ℓ(z) (slides 52–56).
Where we once wanted ℓ(z) ≈ f(z) (with f the target), we now want z + ℓ(z) ≈ f(z), so ℓ(z) = f(z) − z: the layer only has to learn what hasn't been learned yet. If z is already close to the right answer, ℓ(z) barely needs to learn anything. The slides cite the loss-landscape visualizations from Li et al., NeurIPS 2018: with skip connections, the loss surface is much smoother.
From BatchNorm to DyT
- Batch Normalization: compute the mean and standard deviation over a batch, standardize, then scale and shift with learned γ and β; ε in the denominator prevents division by zero.
- Layer Normalization: the original Transformer's choice. It computes the mean and standard deviation over a single vector's own components instead.
- RMSNorm: used by Llama and other LLMs. It divides by the root mean square without subtracting the mean. Slide 62 shows Sebastian Raschka's architecture comparison from GPT-2 to Llama 3.2 in LLMs-from-scratch, noting that by now you can see how newer models differ.
- Dynamic Tanh (DyT): the slide title reads "Meta recently said you don't need normalization!" DyT(x) = γ · tanh(αx) + β, with α initialized to something like 0.5. The paper is Transformers without Normalization.
Normalization formulas (slides 58–64)
LayerNorm: μ = (1/n) Σ x_i, σ = sqrt( (1/n) Σ (x_i − μ)² )
x̂ = (x − μ) / (σ + ε), y = γ x̂ + β (γ, β are learned vectors; numpy-style broadcasting)
RMSNorm: RMS(x) = sqrt( (1/n) Σ x_i² ), RMSNorm(x) = γ · x / RMS(x)
DyT: DyT(x) = γ · tanh(αx) + β
Closing: Transformers aren't just for text
The last two slides (65–66) open the door to the rest of the course. Images, audio, and time series can all go through a Transformer. And the decoder's middle layer, with q from one side and k, v from the other, is a good way to "mix in" information. Take text-to-image AI: a randomly generated "wild idea" serves as Q, a matrix representing the text's meaning produces K and V, and the Transformer's output carries the text's meaning. That thread gets picked up in L11 Text-to-Image.
This week's demo notebook and homework
This lecture has no in-class demo notebook, and the week 5 homework column on the Chang Gung satellite class page says plainly "no homework this week." It's one of two teaching weeks without homework; the other is week 15, "New Trends in Generative AI."
No homework doesn't mean you can skip it. The L04 benchmark assignment is still within its two-week window, and this week is a good time to check your explanations of LLM behavior against L05's ideas.
Self-check
- Why can the row-vector product xA be read as a linear combination of A's rows?
- In one sentence each, what roles do query, key, and value play?
- What problem does dividing by √d_k solve? Would another constant work?
- What attention layers do the encoder and decoder each have? Which one is not self-attention?
- Why does self-attention need positional encoding? In the sin/cos version, which dimensions have the highest frequency?
- Why does the residual z + ℓ(z) make deep networks easier to train?
Something to do tonight: in a Colab, write softmax(Q @ K.T / np.sqrt(d_k)) @ V with numpy, using random matrices for 3 words and d_k = 4. Then remove √d_k and see whether the weights become more winner-take-all. It takes under ten lines to see the effect from slides 32–33 for yourself.
import numpy as np
def softmax(s):
e = np.exp(s - s.max(axis=-1, keepdims=True))
return e / e.sum(axis=-1, keepdims=True)
scores = np.array([3.9, 3.2, 1, 0.3, 1.1])
print(softmax(scores).round(2)) # [0.61 0.3 0.03 0.02 0.04]
print(softmax(scores / np.sqrt(5)).round(2)) # [0.4 0.29 0.11 0.08 0.11]
Further reading
This post stands on its own. For more rigorous or more hands-on versions:
- From positional encoding to causal self-attention: CMU 07-280 Lecture 20
- Attention and Transformers from a deep-learning perspective: CMU 11-785 Lecture 18
- Full courses on Transformers and LLMs: Stanford CME295 guide, Stanford CS224N guide
- Writing a Transformer language model from scratch: Stanford CS336 guide
- Another Mandarin-taught route: NTU Hung-yi Lee ML 2026 guide
Series navigation: Series overview | Previous: L04 LLMs Are Simpler Than You Think | Next: L06 LLM Applications and Ethics
References
- Chang Gung satellite class page: Generative AI 2025 (schedule; no week 5 homework) (in Chinese)
- Lecture 05: The complete guide to Transformers (YouTube recording, 2025-03-18) (in Mandarin)
- 1132 Generative AI recordings playlist (in Mandarin)
- GenAI05 The Mathematics of Transformers slides (Google Drive) (in Chinese)
- 1132 slide folder entry point (yenlung.me/1132GenAI)
- Vaswani et al. 2017: Attention Is All You Need
- He et al. 2015: Deep Residual Learning for Image Recognition (ResNet)
- Li et al. 2018: Visualizing the Loss Landscape of Neural Nets
- Ba et al. 2016: Layer Normalization
- Zhang & Sennrich 2019: Root Mean Square Layer Normalization
- Zhu et al. 2025: Transformers without Normalization (DyT)
- rasbt/LLMs-from-scratch
Loading...