🌏 中文版
This is post 4 of the Reading NTHU Hung-Yu Kao Natural Language Processing series. Post 2 introduced RNN language models, which read one token at a time and predict the next. HW1 tested word vectors. This post takes on a problem RNN language models never faced: how do you design a model when the input and output have different lengths?
The source is the 33-slide deck W3_Sequence-to-sequence Models and Attention Mechanisms.pdf from the IKMLab course repo. The 2025 schedule lists it in the W3 row with two recordings, W3 Tue and W3 Thu (lectures are in Mandarin). That row's Topics column says "Introduction to NLP (Language model)", but it is a syllabus template, so the slides are the source of truth. This post is based on the slides only; I did not check it against the recordings segment by segment.
The problem: translation lengths do not line up
The deck opens by noting that many advances in NLP language models were first driven by translation. It shows a few examples, including the Chinese idiom 來都來了 rendered as "Since we're already here…".
The real point is on slide 3. The same sentence has different lengths in three languages:
| Language | Sentence | Tokens |
|---|---|---|
| English | Look over there | 3 |
| Chinese | 請看那邊 | 4 |
| Japanese | あそこを見てください | 12 |
Classification has a fixed output size, so you can put an FFN on top of an RNN. Translation output may be longer or shorter than the input. The slide's conclusion: the hidden state has to encode the original sequence and pass it on to the generation side.
Encoder-decoder: read everything, then speak
A seq2seq model maps a variable-length input sequence to a variable-length output sequence. The slides list three uses: machine translation, summarization, and dialogue generation.
Generating a sequence with an RNN works like this: feed the input one element at a time, update the hidden state at each step, emit one output element per step, and repeat until a target length or an end-of-sequence token. Generation starts from a special start token.
Translation uses the encoder-decoder architecture (slide 10, drawn from Jurafsky & Martin's SLP3), which has three parts:
- Encoder: reads x₁:ₙ and produces contextual representations h₁:ₙ.
- Context vector: a function of h₁:ₙ that hands the gist of the input to the decoder. The simplest version is the encoder's last hidden state, which is said to "contain all the information of the input."
- Decoder: starts from the context vector and generates an output of any length.
Squeezing a whole sentence into one vector is exactly what attention will later fix.
Vanishing gradients in RNNs
Slides 11–14 walk backpropagation through the shared parameter W. Every time step reuses W, so the earlier an input is, the more factors its gradient gets multiplied by. The slides conclude:
- If those factors are below 1, the gradient decays exponentially: vanishing gradients.
- If they are above 1, it grows quickly: exploding gradients.
Vanishing gradients cause three problems. Important information from early steps is lost as sequences get longer. Parameter updates for early steps slow down or stall. Long generated sequences come out poorly and may become incoherent.
LSTM: gates that decide what to keep and forget
LSTM was proposed to address vanishing gradients in RNNs. Besides the hidden state hₜ, it keeps a cell state cₜ and uses gates to control what flows in and out:
| Component | What the slides say |
|---|---|
| Forget gate | A sigmoid decides which past memory to discard |
| Input gate | A sigmoid decides which new information to write into the cell state |
| Candidate memory | A tanh produces the new content that might be written |
| Output gate | Controls how much of the cell state flows to the hidden state |
Slide 21 explains the four parts with one sentence: "The cat chased the mouse, and then it climbed a tree."
- At "climbed a tree", the forget gate lowers the weight of "chased the mouse" because the focus has moved to a new action.
- At "climbed", the input gate stores the new action in memory.
- The candidate memory encodes "climb + tree".
- When producing "tree", the output gate pulls the location being climbed out of memory.
Gates and memory cells let an LSTM keep or drop information selectively. Gradient flow is steadier, and long-range dependencies are easier to capture than with a plain RNN. The slides say LSTM "partially" avoids the problem.
The two problems left in the RNN family
Slide 23 spells them out:
- Vanishing/exploding gradients: LSTM eases them, but they remain.
- Poor parallelism: each step waits for the previous hidden state. That limits training speed on large datasets and caps how big language models can grow.
Attention: look back at the whole sentence at every step
The core idea of attention: when producing each output, let the model focus on the most relevant parts of the input. The slides present it in two stages.
Stage one: attention with RNNs
Attention first appeared alongside RNNs, in Bahdanau, Cho & Bengio (2014). When the decoder computes a new hidden state sₜ, it scores sₜ against every input token (encoded by a bidirectional RNN) and takes a weighted sum of the inputs. An alignment model produces the scores, measuring how well the inputs around position j match the output at position i.
The decoder no longer depends on one context vector. It sees the whole input at every step. Slide 27 shows the paper's translation alignment plot.
Stage two: drop the RNN
Slide 28 gives two reasons to remove the RNN. Its main job, extracting sequential features, can be done with simpler and cheaper methods. And RNNs cannot be parallelized, which limits scale.
Attention without an RNN multiplies the input by three learned weight matrices to get Q, K, and V, then computes an N×N score matrix QᵀK, where N is the text length.
Why divide by √d
Slides 30–31 explain the scaling factor. Suppose each component of q and k is a random variable with mean 0 and variance 1. Their dot product then has variance d. To bring the variance back to 1, divide the scores by √d.
The summary slide:
- RNN is the basic neural network for NLP tasks; LSTM modifies it to ease gradient problems.
- Attention with RNNs avoids vanishing gradients; attention without RNNs is parallelizable and simpler.
Attention without RNNs is the heart of the Transformer, which this series covers in full in Transformers and Self-Attention.
Things to try after reading
- Check the variance yourself. In NumPy, draw two standard normal vectors with d = 64, take ten thousand dot products, measure the variance, then divide by 8. Ten lines of code confirm slide 31.
- Annotate the LSTM example in another language. Translate "The cat chased the mouse, and then it climbed a tree" into a language you know, mark what each gate should do word by word, and compare with slide 21.
- Warm up for the next post. HW2 has you teach a two-layer LSTM arithmetic. The vanishing and exploding gradients here map to that report's question about why gradient clipping is needed.
Further reading
- Previous in this series: HW1 Word Analogy
- Next in this series: PyTorch TA Session + HW2 Arithmetic as a Language
- The same material in English-language courses: CS224N: RNNs and Language Models, CMU 11-785: Language Models and Translation, CMU 11-785: Attention and Transformers
- Back to the series overview
References
- W3_Sequence-to-sequence Models and Attention Mechanisms.pdf (course slides)
- 2025 schedule README
- Fall 2025 W3 Tue recording (in Mandarin)
- Fall 2025 W3 Thu recording (in Mandarin)
- Bahdanau, Cho & Bengio (2014), Neural Machine Translation by Jointly Learning to Align and Translate
- Jurafsky & Martin, Speech and Language Processing (3rd ed. draft)
- Hochreiter & Schmidhuber (1997), Long Short-Term Memory
Loading...